Train Custom ModelChinese & English

Train Custom Model: LocateAnything

The Train Custom Model: LocateAnything plugin fine-tunes the NVIDIA LocateAnything-3B AI detection model on your own labeled data, so it learns to find the objects specific to your samples. General-purpose models often p

Updated 2026-07-07User manual

LocateAnything 自定义模型训练 插件用户手册

Train Custom Model: LocateAnything - User Manual

Dragonfly Prototype Apps · Train Custom Model: LocateAnything...

版本 Version 1.0 · 2026-07-04


第一部分 中文手册

目录

1. 简介

2. 适用场景

3. 安装与启用

4. 运行环境与首次配置

5. 界面说明

5.1 分组一:Environment & Base Model(环境与基础模型)

5.2 分组二:Dataset(数据集——每个标注切片添加一个样例)

5.3 分组三:LoRA & Training(LoRA 与训练)

5.4 底部按钮与日志

6. 使用步骤

6.1 构建训练集

6.2 训练

6.3 在推理插件中使用新模型

7. 参数说明

8. 输出结果

9. 常见问题与故障排除

10. 注意事项与已知限制

11. 参考资料

1. 简介

Train Custom Model: LocateAnything(训练自定义 LocateAnything 模型)插件用于在您自己的标注数据上微调 NVIDIA LocateAnything-3B AI 检测模型,使其学会识别您样品中特有的目标物体。通用模型对孔隙、裂纹、特殊细胞等专业对象的检测效果往往不理想;本插件让您直接在 Dragonfly 会话中构建一个小型训练集——选择一幅图像 Channel、一个包含标注对象的 MultiROI、一个切片和一个类别名称(如“孔隙”),逐个添加训练样例——然后一键开始训练。训练结果是一个自定义模型文件夹,只需在 LocateAnything(推理)插件的 Model folder 中选中它,即可在新图像中检测您的目标物体。

底层技术:训练采用 LoRA(低秩适配)参数高效微调,通过 Hugging Face PEFT 库对 LocateAnything-3B 的语言模型部分注入 LoRA 适配层,在自回归目标串上做标准监督微调(SFT),使用 fp16 混合精度与梯度检查点以节省显存。训练完成后 LoRA 权重默认合并回基础模型(merge_and_unload),保存为一个可直接由推理插件加载的完整模型文件夹。整个训练过程在 LocateAnything(推理)插件已搭建的独立 Python 虚拟环境(torch + transformers + peft + accelerate)中以子进程方式运行,完全在 Dragonfly 之外执行,不会触碰 Dragonfly 自带的 Python。

许可证要点:LocateAnything-3B 基础模型来自 NVIDIA,训练所得模型继承其许可条款,仅限非商业 / 科研用途。请勿将微调模型用于商业目的。

实验性功能:NVIDIA 未公布 LocateAnything 的官方训练方案,本插件采用标准 LoRA SFT 作为近似实现。训练完成后,在推理插件中请把生成模式(generation_mode)设为 slow 或 hybrid 再运行检测,以匹配训练目标。

2. 适用场景

本插件适用于通用模型检测效果不佳的目标对象。只需少量标注切片即可让模型学会您的特定对象类别,典型场景包括:

  • 材料科学与工业 CT 扫描:孔隙、裂纹、夹杂物、纤维等缺陷或微结构的自动检测;
  • 生命科学图像:细胞及其他生物结构的定位与计数;
  • 任何已经用 Dragonfly MultiROI 完成标注、希望把标注经验“教”给 AI 检测模型的数据集;
  • LoRA 微调所需的数据量小——一个由若干标注切片构成的精选小数据集即可开始训练。

工作流定位:本插件负责“训练”,检测(推理)由配套的 LocateAnything 推理插件完成。两者共用同一虚拟环境和同一基础模型,训练输出的模型文件夹可直接被推理插件加载使用。

3. 安装与启用

本插件随 Prototype Labs Full Package(完整安装包)一起分发,安装步骤如下:

1. 将安装包 zip 解压到任意较短路径(如 C:\PL\,避免过深的目录导致 Windows 260 字符路径限制);

2. 双击 `Install_FullPackage.bat` 运行安装器;

3. 在弹出的对话框中勾选想要的 Prototype Apps。注意:本插件默认未勾选(所有重型插件默认关闭),请手动勾选 “Train Custom Model: LocateAnything...”;建议同时勾选它依赖的 “LocateAnything...”推理插件;

4. 点击 Install,等待控制台完成;

5. 完全退出并重启 Dragonfly(菜单只在启动时扫描)。

重启后,菜单项出现在 Prototype Apps ▸ Train Custom Model: LocateAnything...(归属 “Train Custom Model”分组)。点击即打开名为 Train Custom LocateAnything 的浮动面板窗口。

以后启用 / 停用:打开 Developer ▸ Prototype Labs... ▸ Menu Item Manager,在底部 “Prototype Apps (Full Package)”列表中勾选或取消勾选本插件,重启 Dragonfly 生效。停用不会删除任何已有设置或环境,重新启用立即可用。也可以随时重跑安装器(上次的选择即默认值):%LOCALAPPDATA%\DragonflyPrototypeLabs\FullPackage\installer\Install_FullPackage.bat。

卸载:双击 `Uninstall_FullPackage.bat`(同样保存在上述 installer 目录)。它移除所有 Full Package 菜单项、插件和中央存储,但保留各插件已搭建的环境;结束时会列出可手动删除的路径。安装全部位于当前用户目录(%LOCALAPPDATA%),无需管理员权限。

4. 运行环境与首次配置

本插件自身不下载、不安装任何东西,也没有独立的 Setup Environment 步骤——它直接复用 LocateAnything(推理)插件已搭建的虚拟环境(其中已含 torch、transformers、peft、accelerate)。因此首次使用前必须先完成以下前置条件:

1. 安装 LocateAnything(推理)插件 并在其面板中运行一次 Setup Environment(该步骤用 Dragonfly 自带 Python 建立 venv 并安装 CUDA 版 torch 和 transformers,约 5–15 分钟、下载数 GB,需联网);

2. 在推理插件中设置好 Model folder(LocateAnything-3B 基础模型权重,需通过 huggingface-cli 单独下载到 Dragonfly 之外的用户目录,无需令牌);

3. 确认本机有 NVIDIA GPU 且显存不低于 24 GB(在 3B 视觉语言模型上做 LoRA 微调的最低要求;48 GB 显存则非常充裕)。Turing 架构显卡不支持 bf16,训练自动使用 fp16。

完成上述准备后打开本插件面板,基础模型文件夹与推理 venv python 两栏会自动从推理插件保存的配置中预填(读取推理插件安装目录下的 locateanything_config.json 以及 %LOCALAPPDATA%\LocateAnything\config.json),通常无需手动修改。

本插件涉及的磁盘路径:

  • %LOCALAPPDATA%\LocateAnything\config.json —— 推理插件的共享配置(venv python 路径、基础模型路径、运行模式),本插件读取它作为默认值;
  • %LOCALAPPDATA%\LocateAnythingTrain\config.json —— 本插件自己保存的面板设置(每次添加样例 / 训练 / 关闭面板时自动写入);
  • C:\TrainJobs —— 默认作业根目录(可在面板中修改),数据集文件夹和每次训练的作业文件夹都创建在这里。

联网与 GPU:本插件运行本身不需要联网(前提环境已就绪),但必须有 NVIDIA GPU——训练脚本启动时会检查 CUDA,不可用即报错。高级选项:运行模式可选 wsl,通过用户自行配置的 WSL 发行版在 Linux 侧运行训练脚本(路径自动转换为 /mnt/... 形式);一般 Windows 用户保持默认 windows 即可。

如果面板提示推理 venv 未设置或基础模型缺失,请回到 LocateAnything(推理)插件先完成 Setup Environment 和模型下载——这是本插件唯一的“环境搭建”途径。

5. 界面说明

面板为单页布局:顶部是一段蓝色说明文字(概述 LoRA 微调流程与实验性提示),中间是可滚动的三个分组框,底部是 Train / Open Output Folder 两个按钮、结果状态行和只读日志窗口。以下按分组逐一说明每个控件。

5.1 分组一:Environment & Base Model(环境与基础模型)

  • Base model folder(文本框 + “…”浏览按钮):LocateAnything-3B 基础模型文件夹。默认自动填入推理插件的 Model folder;
  • Inference venv python(文本框):推理插件 venv 中 python 解释器的完整路径,来自推理插件的 Setup Environment,自动预填;
  • Run mode(下拉框,选项 windows / wsl,默认 windows):训练子进程的运行方式;wsl 为高级选项,需自行配置好 WSL 发行版;
  • Job root(文本框,默认 C:\TrainJobs):作业根目录,数据集与训练作业文件夹均创建于此。

5.2 分组二:Dataset(数据集——每个标注切片添加一个样例)

  • Image Channel(下拉框 + Refresh 按钮):选择当前 Dragonfly 会话中的图像 Channel;Refresh 重新扫描会话中的 Channel 和 MultiROI;
  • MultiROI (labels)(下拉框):选择包含标注的 MultiROI——其中每个 label(标签)对应一个目标物体,导出时被转换为一个边界框;
  • Orientation(下拉框,XY / XZ / YZ,默认 XY):取切片的方向;
  • Slice index(数字框):切片序号,范围随所选 Channel 与方向自动更新,初次默认取中间切片;
  • Category (prompt)(文本框,默认 objects):本 MultiROI 所有标签共同的类别名(即训练提示词),例如 pores;
  • Max image size (px)(数字框,256–4096,步长 128,默认 1024):导出图像的最长边上限,超出则等比降采样;
  • Intensity window(lo / hi 两个文本框,留空 = auto):灰度窗宽;留空时自动取该切片的 1%–99% 分位数拉伸到 8 位;
  • Add example 按钮:把当前 Channel + MultiROI + 切片 + 类别导出为一个训练样例(追加进数据集);若尚无数据集会自动新建一个;
  • New dataset 按钮:在 Job root 下新建一个空数据集文件夹(dataset_日期_时间),重新开始收集样例;
  • Dataset 状态行:显示当前数据集的样例数量与文件夹路径。

5.3 分组三:LoRA & Training(LoRA 与训练)

  • Output model folder(文本框 + “…”浏览按钮):训练输出的模型文件夹,例如 C:\models\LocateAnything-3B-mine;
  • LoRA rank (r)(数字框,1–256,默认 16):LoRA 低秩矩阵的秩;
  • LoRA alpha(数字框,1–512,默认 32):LoRA 缩放系数;
  • LoRA dropout(文本框,默认 0.05):LoRA 层的 dropout 比例;
  • Learning rate(文本框,默认 0.0001):学习率;
  • Epochs(数字框,1–200,默认 5):训练轮数;
  • Grad accumulation(数字框,1–64,默认 8):梯度累积步数(实际 batch size 固定为 1,用累积模拟更大批量);
  • Gradient checkpointing (save VRAM)(复选框,默认勾选):梯度检查点,以少量计算换取显存节省;
  • Merge LoRA into a full model (drop-in for inference)(复选框,默认勾选):训练后把 LoRA 合并进基础模型并保存完整模型;取消勾选则仅保存 LoRA 适配器。

5.4 底部按钮与日志

  • Train(蓝色按钮):校验设置后启动训练子进程;
  • Open Output Folder:在资源管理器中打开输出模型文件夹(或 Job root);
  • 结果状态行:显示 Training… / Done / Failed 及提示信息;
  • 日志窗口(只读):滚动显示样例导出信息、训练进度百分比、epoch / step / loss 以及错误详情。

6. 使用步骤

完整工作流分三个阶段:构建训练集 → 训练 → 在推理插件中使用新模型。

6.1 构建训练集

1. 在 Dragonfly 中加载图像 Channel,并准备好标注:一个 MultiROI,其中每个 label 是同一类别的一个目标物体(例如每个孔隙一个 label);

2. 打开 Prototype Apps ▸ Train Custom Model: LocateAnything...,点 Refresh,日志会显示找到的 Channel 与 MultiROI 数量;

3. 在下拉框中选好 Image Channel 与 MultiROI (labels),选择 Orientation 与 Slice index(选一个有标注的切片);

4. 在 Category (prompt) 中输入类别名,如 pores(英文提示词效果更稳定);

5. (可选)调整 Max image size 与 Intensity window;

6. 点 Add example。插件在 UI 线程读取该切片,把图像窗化为 8 位 RGB 保存成 input_<k>.npy,把 MultiROI 同一切片上的每个 label 转成边界框(归一化到 [0,1000] 网格)并生成监督目标串,追加一行到 dataset.jsonl。日志显示“Added example #n: m box(es), category 'xxx'”;

7. 换切片(或换 Channel / MultiROI)重复第 3–6 步,为每个标注切片各添加一个样例。数据集状态行实时显示样例总数;

8. 如需重新开始,点 New dataset 新建一个空数据集文件夹。

6.2 训练

1. 确认 Base model folder 与 Inference venv python 已正确填入(通常自动预填);

2. 设置 Output model folder(新模型的保存位置);

3. 按需调整 LoRA 参数(rank / alpha / dropout / learning rate / epochs / grad accumulation),首次训练建议保持默认;

4. 点 Train。插件在作业目录写入 train_config.json,然后用推理插件的 venv python 启动 train_runner.py 子进程;

5. 观察日志:加载基础模型 → 注入 LoRA → 逐 epoch 训练(每步显示进度百分比、step 与 loss)→ 合并 LoRA 并保存模型;

6. 完成后状态行显示 Done,日志给出输出文件夹路径;点 Open Output Folder 可直接打开。

训练在后台子进程中运行,期间可以继续使用 Dragonfly;关闭面板会终止正在进行的训练。训练时长取决于样例数、epochs 与 GPU 性能。

6.3 在推理插件中使用新模型

1. 打开 LocateAnything(推理)插件;

2. 把它的 Model folder 指向本插件的输出模型文件夹;

3. 把生成模式(generation_mode)设为 slow 或 hybrid(LoRA SFT 训练的是自回归目标,纯 fast/MTP 模式可能无法体现微调效果);

4. 在新图像上用您训练时的类别名作为提示词运行检测。

7. 参数说明

参数

默认值

说明

Base model folder

(推理插件的 Model folder)

LocateAnything-3B 基础模型文件夹;自动从推理插件配置预填

Inference venv python

(推理插件的 venv)

训练使用的 Python 解释器;由推理插件的 Setup Environment 创建

Run mode

windows

训练子进程运行方式:windows(本机)或 wsl(用户自配的 WSL 发行版)

Job root

C:\TrainJobs

数据集与训练作业文件夹的根目录

Orientation

XY

取切片方向:XY / XZ / YZ

Slice index

中间切片

切片序号;范围随 Channel 与方向自动更新

Category (prompt)

objects

该 MultiROI 全部标签的类别名,即训练与检测的提示词

Max image size (px)

1024

导出图像最长边上限(256–4096,步长 128),超出则等比降采样

Intensity window lo/hi

auto

灰度窗;留空自动取切片 1%–99% 分位数

Output model folder

(空)

训练输出的模型文件夹,必填

LoRA rank (r)

16

LoRA 低秩矩阵的秩(1–256);越大可学容量越高、显存开销越大

LoRA alpha

32

LoRA 缩放系数(1–512)

LoRA dropout

0.05

LoRA 层 dropout 比例

Learning rate

0.0001

AdamW 优化器学习率

Epochs

5

训练轮数(1–200)

Grad accumulation

8

梯度累积步数(1–64);batch size 固定为 1

Gradient checkpointing

勾选

梯度检查点,节省显存

Merge LoRA into a full model

勾选

训练后合并 LoRA 保存完整模型;取消则仅保存适配器

8. 输出结果

本插件不在 Dragonfly 中创建新对象(不生成 Channel / ROI / Mesh),全部输出为磁盘上的文件夹与文件:

  • 数据集文件夹 Job root\dataset_日期_时间\:每个样例一个 input_<k>.npy(8 位 RGB 切片图像)+ 一个 dataset.jsonl(每行一个样例:图像路径、类别提示词、[0,1000] 归一化边界框和监督目标串);
  • 训练作业文件夹 Job root\train_日期_时间\:train_config.json(本次训练的完整配置)、status.json(实时进度)、train_result.json(最终结果摘要:成功与否、样例数、epochs、耗时等);
  • 输出模型文件夹(您指定的 Output model folder):合并后的完整模型(safetensors 权重 + tokenizer / processor / chat template 及自定义代码文件,可被推理插件直接加载),另含 lora_adapter\ 子文件夹保存原始 LoRA 适配器以供参考。

查看方式:训练完成后点 Open Output Folder 打开输出文件夹;检测效果则在 LocateAnything(推理)插件中把 Model folder 指向该文件夹后运行验证。

9. 常见问题与故障排除

问题 1:点 Train 提示 “Inference venv not set (run Setup Environment in the inference plugin)”。
原因:尚未安装 LocateAnything(推理)插件或未运行其 Setup Environment。解决:先在推理插件面板中完成 Setup Environment;完成后本面板的 Inference venv python 会自动预填(或手动填入该 venv 中 python.exe 的完整路径)。

问题 2:提示 “Base model folder not set/missing”。
原因:基础模型路径为空或文件夹不存在。解决:该栏默认取推理插件的 Model folder——请先按推理插件手册用 huggingface-cli 下载 LocateAnything-3B 权重并在推理插件中设置好 Model folder,再回到本面板(或用“…”按钮手动选择模型文件夹)。

问题 3:Add example 后日志显示 0 box(es)。
原因:所选切片上该 MultiROI 没有任何标签像素。解决:换一个确实包含标注的 Slice index / Orientation;注意每个 label 在该切片上的像素范围决定其边界框,跨切片的标注只统计当前切片。

问题 4:训练报 CUDA out of memory(显存不足)。
解决:保持 Gradient checkpointing 勾选;调小 Max image size(如 768 或 512)后重新添加样例再训练;降低 LoRA rank。硬性要求是显存不低于 24 GB 的 NVIDIA GPU——显存更小的显卡不受支持。

问题 5:训练报 “CUDA not available to PyTorch”。
原因:推理插件 venv 中的 torch 检测不到 GPU(驱动缺失或装成了 CPU 版)。解决:确认 NVIDIA 驱动正常(nvidia-smi 可用),必要时在推理插件中重新运行 Setup Environment 安装 CUDA 版 torch。

问题 6:Channel / MultiROI 下拉框是空的。
解决:确认数据已在当前 Dragonfly 会话中加载,然后点 Refresh。

问题 7:训练完成的模型在推理插件里检测不到微调的对象。
解决:确认推理插件的 generation_mode 设为 slow 或 hybrid(不要用纯 fast 模式);确认提示词与训练时的 Category 一致;样例过少或标注质量不高时适当增加标注切片和 epochs。

问题 8:报 “peft not installed in the venv”。
原因:推理插件 venv 缺少 peft / accelerate 包。解决:在推理插件中重新运行 Setup Environment(其依赖清单已包含 peft 与 accelerate;本插件的 runner\requirements.txt 仅作参考,正常情况下无需手动安装)。

10. 注意事项与已知限制

  • 实验性功能:NVIDIA 未公布官方训练方案,本插件的 LoRA SFT 是近似实现;训练质量高度依赖标注数据的数量与质量;
  • 许可限制:微调模型继承 NVIDIA LocateAnything 许可,仅限非商业 / 科研用途;
  • 每个 MultiROI 只对应一个类别:一次添加样例时,该 MultiROI 的所有标签都视为您输入的同一类别;按标签区分多类别暂不支持;
  • 仅支持 detect(检测)任务:点定位、grounding 等任务类型以及从磁盘导入现成数据集暂不支持;
  • 训练精度为 fp16(Turing 架构 GPU 无 bf16),batch size 固定为 1,通过梯度累积模拟更大批量;
  • 样例是 2D 切片:训练集由单张切片图像 + 该切片上的标签边界框构成;MultiROI 应与 Channel 几何对齐(切片尺寸不一致时插件用最近邻缩放对齐掩膜,可能引入少量偏差);
  • 关闭面板会终止训练:请等训练完成(状态行显示 Done)再关闭;
  • 菜单变更需重启:启用 / 停用本插件后需重启 Dragonfly 才能生效。

11. 参考资料

  • NVIDIA LocateAnything-3B:本插件微调的基础视觉语言检测模型(模型权重经 Hugging Face 的 huggingface-cli 下载,详见推理插件手册);
  • Hugging Face Transformers:模型与处理器加载(trust_remote_code 方式);
  • Hugging Face PEFT(peft ≥ 0.11):LoRA 参数高效微调实现;
  • Hugging Face Accelerate(accelerate ≥ 0.34):训练加速支持库;
  • 配套插件:LocateAnything(推理)插件 用户手册——环境搭建(Setup Environment)、基础模型下载与检测运行均在该手册中详述。


Part II English Manual

Contents

1. Overview

2. Use Cases

3. Installation and Enabling

4. Runtime Environment and First-Time Setup

5. User Interface

5.1 Group 1: Environment & Base Model (from the inference plugin)

5.2 Group 2: Dataset (add one example per labeled slice)

5.3 Group 3: LoRA & Training

5.4 Bottom Buttons and Log

6. Step-by-Step Usage

6.1 Build the Training Set

6.2 Train

6.3 Use the New Model in the Inference Plugin

7. Parameter Reference

8. Outputs

9. FAQ and Troubleshooting

10. Notes and Known Limitations

11. References

1. Overview

The Train Custom Model: LocateAnything plugin fine-tunes the NVIDIA LocateAnything-3B AI detection model on your own labeled data, so it learns to find the objects specific to your samples. General-purpose models often perform poorly on specialized targets such as pores, cracks, or particular cell types. With this plugin you build a small training set directly from your Dragonfly session — pick an image Channel, a MultiROI containing your labeled objects, a slice, and a category name (e.g. "pores"), then add examples one by one — and start training with a single click. The result is a custom model folder that you simply select as the Model folder in the LocateAnything (inference) plugin to detect your objects in new images.

Under the hood, training uses LoRA (low-rank adaptation) parameter-efficient fine-tuning: the Hugging Face PEFT library injects LoRA adapter layers into the language-model part of LocateAnything-3B, and the plugin runs standard supervised fine-tuning (SFT) on the autoregressive target string, using fp16 mixed precision and gradient checkpointing to save VRAM. After training, the LoRA weights are by default merged back into the base model (merge_and_unload) and saved as a full model folder that the inference plugin can load directly. The whole training process runs as a subprocess in the dedicated Python virtual environment already built by the LocateAnything (inference) plugin (torch + transformers + peft + accelerate) — entirely outside Dragonfly, never touching Dragonfly's own Python.

License: the LocateAnything-3B base model comes from NVIDIA; any fine-tuned model inherits its license terms and is for non-commercial / research use only. Do not use trained models commercially.

Experimental feature: NVIDIA has not published an official training recipe for LocateAnything; this plugin implements standard LoRA SFT as an approximation. After training, set the generation mode (generation_mode) in the inference plugin to slow or hybrid before running detection, to match what was trained.

2. Use Cases

This plugin is for anyone whose objects of interest are not detected well by the general-purpose model. A handful of labeled slices is enough to teach the model your specific object category. Typical scenarios:

  • Materials science and industrial CT scans: automatic detection of pores, cracks, inclusions, fibers, and other defects or microstructures;
  • Life-science images: locating and counting cells and other biological structures;
  • Any dataset you have already annotated with Dragonfly MultiROIs, where you want to "teach" your labeling expertise to an AI detection model;
  • LoRA fine-tuning needs little data — a curated small dataset of a few labeled slices is enough to start training.

Workflow positioning: this plugin does the training; detection (inference) is done by the companion LocateAnything inference plugin. The two share the same virtual environment and the same base model, and the trained model folder produced here can be loaded directly by the inference plugin.

3. Installation and Enabling

This plugin ships with the Prototype Labs Full Package installer:

1. Unzip the package to a short path (e.g. C:\PL\; avoid deep folders because of the Windows 260-character path limit);

2. Double-click `Install_FullPackage.bat`;

3. In the dialog, tick the Prototype Apps you want. Note: this plugin is unchecked by default (all heavy plugins are off by default) — tick "Train Custom Model: LocateAnything..." manually, and it is recommended to also tick the "LocateAnything..." inference plugin it depends on;

4. Click Install and wait for the console to finish;

5. Quit Dragonfly completely and restart it (menus are only scanned at startup).

After the restart, the menu entry appears under Prototype Apps ▸ Train Custom Model: LocateAnything... (in the "Train Custom Model" group). Clicking it opens a floating panel window titled Train Custom LocateAnything.

Enable / disable later: open Developer ▸ Prototype Labs... ▸ Menu Item Manager — the "Prototype Apps (Full Package)" list at the bottom has a checkbox per app; tick or untick this plugin and restart Dragonfly to apply. Disabling never deletes any settings or environments; re-enabling is instant. You can also re-run the installer anytime (your previous choices are the new defaults): %LOCALAPPDATA%\DragonflyPrototypeLabs\FullPackage\installer\Install_FullPackage.bat.

Uninstall: double-click `Uninstall_FullPackage.bat` (also kept in the installer folder above). It removes all Full-Package menu items, plugins and the central store, but keeps every plugin environment; the removable paths are listed at the end. Everything installs per-user (%LOCALAPPDATA%); no admin rights are needed.

4. Runtime Environment and First-Time Setup

This plugin downloads and installs nothing of its own and has no separate Setup Environment step — it directly reuses the virtual environment already built by the LocateAnything (inference) plugin (which contains torch, transformers, peft and accelerate). Before first use you must therefore complete these prerequisites:

1. Install the LocateAnything (inference) plugin and run its Setup Environment once (it builds a venv from Dragonfly's bundled Python and pip-installs CUDA torch + transformers; about 5–15 minutes and a few GB of downloads; requires internet);

2. Set the inference plugin's Model folder (the LocateAnything-3B base weights, downloaded separately via huggingface-cli — no token needed — to a user path outside Dragonfly);

3. Have an NVIDIA GPU with at least 24 GB of VRAM (the minimum for LoRA on a 3B vision-language model; 48 GB is ample). Turing GPUs have no bf16, so training automatically uses fp16.

Once these are done, open this plugin's panel: the Base model folder and Inference venv python fields are pre-filled automatically from the inference plugin's saved configuration (read from locateanything_config.json in the inference plugin's install folder and from %LOCALAPPDATA%\LocateAnything\config.json). Normally no manual editing is needed.

Disk paths used by this plugin:

  • %LOCALAPPDATA%\LocateAnything\config.json — the inference plugin's shared configuration (venv python path, base model path, run mode); this plugin reads it for defaults;
  • %LOCALAPPDATA%\LocateAnythingTrain\config.json — this plugin's own saved panel settings (written automatically on each Add example / Train / panel close);
  • C:\TrainJobs — the default job root (configurable in the panel); dataset folders and per-run job folders are created here.

Internet and GPU: running this plugin itself needs no internet (once the prerequisites are in place), but an NVIDIA GPU is mandatory — the training script checks for CUDA at startup and fails if it is unavailable. Advanced option: the run mode can be set to wsl to run the training script inside a user-configured WSL distribution (paths are translated to /mnt/... form automatically); typical Windows users should keep the default windows.

If the panel reports that the inference venv is not set or the base model is missing, go back to the LocateAnything (inference) plugin and complete its Setup Environment and model download first — that is the only "environment setup" path for this plugin.

5. User Interface

The panel is a single page: a blue introductory note at the top (summarizing the LoRA workflow and the experimental caveat), three group boxes in a scrollable area, and at the bottom the Train / Open Output Folder buttons, a result status line, and a read-only log window. Each group is described below, control by control.

5.1 Group 1: Environment & Base Model (from the inference plugin)

  • Base model folder (text field + "…" browse button): the LocateAnything-3B base model folder. Defaults to the inference plugin's Model folder automatically;
  • Inference venv python (text field): the full path of the python interpreter inside the inference plugin's venv, created by its Setup Environment; pre-filled automatically;
  • Run mode (dropdown, windows / wsl, default windows): how the training subprocess is launched; wsl is an advanced option requiring a user-configured WSL distribution;
  • Job root (text field, default C:\TrainJobs): the root folder where dataset and training-job folders are created.

5.2 Group 2: Dataset (add one example per labeled slice)

  • Image Channel (dropdown + Refresh button): pick an image Channel from the current Dragonfly session; Refresh rescans the session for Channels and MultiROIs;
  • MultiROI (labels) (dropdown): pick the MultiROI holding your annotations — each label is one object instance; at export each label becomes one bounding box;
  • Orientation (dropdown, XY / XZ / YZ, default XY): the slicing direction;
  • Slice index (spinbox): the slice number; the range updates automatically with the selected Channel and orientation, and initially defaults to the middle slice;
  • Category (prompt) (text field, default objects): the category name shared by all labels of this MultiROI — i.e. the training prompt, e.g. pores;
  • Max image size (px) (spinbox, 256–4096, step 128, default 1024): the maximum long-edge size of the exported image; larger slices are downscaled proportionally;
  • Intensity window (lo / hi text fields, empty = auto): the grayscale window; when empty, the 1st–99th percentile of the slice is stretched to 8-bit automatically;
  • Add example button: exports the current Channel + MultiROI + slice + category as one training example (appended to the dataset); a new dataset is created automatically if none exists yet;
  • New dataset button: creates a fresh empty dataset folder (dataset_<date>_<time>) under the Job root to start collecting examples anew;
  • Dataset status line: shows the current example count and the dataset folder path.

5.3 Group 3: LoRA & Training

  • Output model folder (text field + "…" browse button): where the trained model is saved, e.g. C:\models\LocateAnything-3B-mine;
  • LoRA rank (r) (spinbox, 1–256, default 16): the rank of the LoRA low-rank matrices;
  • LoRA alpha (spinbox, 1–512, default 32): the LoRA scaling factor;
  • LoRA dropout (text field, default 0.05): dropout ratio of the LoRA layers;
  • Learning rate (text field, default 0.0001): the learning rate;
  • Epochs (spinbox, 1–200, default 5): number of training epochs;
  • Grad accumulation (spinbox, 1–64, default 8): gradient-accumulation steps (the actual batch size is fixed at 1; accumulation simulates a larger batch);
  • Gradient checkpointing (save VRAM) (checkbox, checked by default): trades a little compute for VRAM savings;
  • Merge LoRA into a full model (drop-in for inference) (checkbox, checked by default): merges the LoRA weights into the base model after training and saves a full model; if unchecked, only the LoRA adapter is saved.

5.4 Bottom Buttons and Log

  • Train (blue button): validates the settings and launches the training subprocess;
  • Open Output Folder: opens the output model folder (or the Job root) in Explorer;
  • Result status line: shows Training… / Done / Failed with a message;
  • Log window (read-only): streams example-export messages, training progress percentage, epoch / step / loss values, and error details.

6. Step-by-Step Usage

The full workflow has three stages: build the training set, train, then use the new model in the inference plugin.

6.1 Build the Training Set

1. Load your image Channel in Dragonfly and prepare the annotations: a MultiROI in which each label is one object instance of the same category (e.g. one label per pore);

2. Open Prototype Apps ▸ Train Custom Model: LocateAnything... and click Refresh; the log reports how many Channels and MultiROIs were found;

3. Select the Image Channel and the MultiROI (labels), choose the Orientation and a Slice index that actually carries labels;

4. Type the category name in Category (prompt), e.g. pores (English prompts are the most reliable);

5. (Optional) adjust Max image size and the Intensity window;

6. Click Add example. The plugin reads that slice, windows the image to 8-bit RGB and saves it as input_<k>.npy, converts every MultiROI label on the same slice into a bounding box (normalized to the [0,1000] grid), builds the supervised target string, and appends one line to dataset.jsonl. The log shows "Added example #n: m box(es), category 'xxx'";

7. Repeat steps 3–6 for every labeled slice (you may switch slices, Channels or MultiROIs between adds). The dataset status line shows the running total;

8. To start over, click New dataset to create a fresh empty dataset folder.

6.2 Train

1. Verify that Base model folder and Inference venv python are filled in (normally pre-filled automatically);

2. Set the Output model folder (where the new model will be saved);

3. Adjust the LoRA parameters (rank / alpha / dropout / learning rate / epochs / grad accumulation) as needed — the defaults are a good starting point;

4. Click Train. The plugin writes train_config.json into a job folder and launches train_runner.py as a subprocess using the inference plugin's venv python;

5. Watch the log: base model loading → LoRA injection → epoch-by-epoch training (each step reports progress percentage, step and loss) → LoRA merge and model saving;

6. When finished, the status line shows Done and the log prints the output folder path; click Open Output Folder to open it directly.

Training runs in a background subprocess, so you can keep using Dragonfly meanwhile; closing the panel terminates a training in progress. Training time depends on the number of examples, epochs, and your GPU.

6.3 Use the New Model in the Inference Plugin

1. Open the LocateAnything (inference) plugin;

2. Point its Model folder at this plugin's output model folder;

3. Set the generation mode (generation_mode) to slow or hybrid (LoRA SFT trains the autoregressive target; the pure fast/MTP mode may not reflect the fine-tune);

4. Run detection on new images using the same category name you trained with as the prompt.

7. Parameter Reference

Parameter

Default

Description

Base model folder

(inference plugin's Model folder)

The LocateAnything-3B base model folder; pre-filled from the inference plugin's config

Inference venv python

(inference plugin's venv)

The Python interpreter used for training; created by the inference plugin's Setup Environment

Run mode

windows

How the training subprocess runs: windows (native) or wsl (user-configured WSL distribution)

Job root

C:\TrainJobs

Root folder for dataset and training-job folders

Orientation

XY

Slicing direction: XY / XZ / YZ

Slice index

middle slice

Slice number; range follows the selected Channel and orientation

Category (prompt)

objects

Category name for all labels of the MultiROI; the training and detection prompt

Max image size (px)

1024

Long-edge cap of the exported image (256–4096, step 128); larger slices are downscaled

Intensity window lo/hi

auto

Grayscale window; empty = automatic 1st–99th percentile of the slice

Output model folder

(empty)

Where the trained model is saved; required

LoRA rank (r)

16

Rank of the LoRA matrices (1–256); higher = more capacity, more VRAM

LoRA alpha

32

LoRA scaling factor (1–512)

LoRA dropout

0.05

Dropout ratio of the LoRA layers

Learning rate

0.0001

AdamW optimizer learning rate

Epochs

5

Number of training epochs (1–200)

Grad accumulation

8

Gradient-accumulation steps (1–64); batch size is fixed at 1

Gradient checkpointing

checked

Gradient checkpointing to save VRAM

Merge LoRA into a full model

checked

Merge LoRA into the base model after training; unchecked saves only the adapter

8. Outputs

This plugin creates no new objects inside Dragonfly (no Channel / ROI / Mesh); all outputs are folders and files on disk:

  • Dataset folder Job root\dataset_<date>_<time>\: one input_<k>.npy per example (the 8-bit RGB slice image) plus dataset.jsonl (one line per example: image path, category prompt, [0,1000]-normalized bounding boxes, and the supervised target string);
  • Training job folder Job root\train_<date>_<time>\: train_config.json (the full configuration of this run), status.json (live progress), and train_result.json (the final summary: success flag, example count, epochs, timing, etc.);
  • Output model folder (the Output model folder you chose): the merged full model (safetensors weights + tokenizer / processor / chat-template and custom code files, loadable directly by the inference plugin), plus a lora_adapter\ subfolder keeping the raw LoRA adapter for reference.

To inspect the results, click Open Output Folder after training; to verify detection quality, point the LocateAnything (inference) plugin's Model folder at the output folder and run it on your data.

9. FAQ and Troubleshooting

Q1: Clicking Train reports "Inference venv not set (run Setup Environment in the inference plugin)".
Cause: the LocateAnything (inference) plugin is not installed, or its Setup Environment has not been run. Fix: complete Setup Environment in the inference plugin's panel first; the Inference venv python field here is then pre-filled automatically (or type the full path of python.exe inside that venv manually).

Q2: "Base model folder not set/missing".
Cause: the base model path is empty or the folder does not exist. Fix: this field defaults to the inference plugin's Model folder — first download the LocateAnything-3B weights via huggingface-cli as described in the inference plugin's manual and set its Model folder, then return here (or pick the model folder manually with the "…" button).

Q3: After Add example the log shows 0 box(es).
Cause: the chosen slice contains no label pixels of that MultiROI. Fix: pick a Slice index / Orientation that actually carries annotations; each label's pixel extent on the current slice defines its bounding box — labels on other slices are not counted.

Q4: Training fails with CUDA out of memory.
Fix: keep Gradient checkpointing enabled; reduce Max image size (e.g. 768 or 512) and re-add the examples before training again; lower the LoRA rank. The hard requirement is an NVIDIA GPU with at least 24 GB of VRAM — smaller GPUs are not supported.

Q5: Training fails with "CUDA not available to PyTorch".
Cause: torch inside the inference plugin's venv cannot see the GPU (missing driver, or a CPU-only build). Fix: verify the NVIDIA driver works (nvidia-smi), and if needed re-run Setup Environment in the inference plugin to install the CUDA build of torch.

Q6: The Channel / MultiROI dropdowns are empty.
Fix: make sure the data is loaded in the current Dragonfly session, then click Refresh.

Q7: The trained model does not detect the fine-tuned objects in the inference plugin.
Fix: make sure the inference plugin's generation_mode is set to slow or hybrid (not pure fast); use exactly the same category name you trained with as the prompt; with very few or low-quality examples, add more labeled slices and/or increase the epochs.

Q8: "peft not installed in the venv".
Cause: the inference plugin's venv lacks the peft / accelerate packages. Fix: re-run Setup Environment in the inference plugin (its dependency list already includes peft and accelerate; this plugin's runner\requirements.txt is for reference only and normally nothing needs to be installed manually).

10. Notes and Known Limitations

  • Experimental: NVIDIA has published no official training recipe; the LoRA SFT implemented here is an approximation, and training quality depends heavily on the quantity and quality of your labeled data;
  • License: fine-tuned models inherit the NVIDIA LocateAnything license — non-commercial / research use only;
  • One category per MultiROI: when adding an example, all labels of the MultiROI are treated as the single category you typed; per-label multi-class training is not yet supported;
  • Detect task only: point localization, grounding tasks, and importing pre-built datasets from disk are not yet supported;
  • fp16 training (Turing GPUs have no bf16); batch size is fixed at 1, with gradient accumulation simulating larger batches;
  • Examples are 2D slices: the training set consists of single slice images plus the label bounding boxes on that slice; the MultiROI should be geometrically aligned with the Channel (if the plane sizes differ, the mask is aligned by nearest-neighbor resizing, which may introduce small offsets);
  • Closing the panel terminates training: wait until the status line shows Done before closing;
  • Menu changes need a restart: enabling / disabling this plugin takes effect after restarting Dragonfly.

11. References

  • NVIDIA LocateAnything-3B: the base vision-language detection model fine-tuned by this plugin (weights downloaded via Hugging Face's huggingface-cli; see the inference plugin's manual);
  • Hugging Face Transformers: model and processor loading (trust_remote_code);
  • Hugging Face PEFT (peft >= 0.11): the LoRA parameter-efficient fine-tuning implementation;
  • Hugging Face Accelerate (accelerate >= 0.34): training acceleration support library;
  • Companion plugin: the LocateAnything (inference) plugin user manual — environment setup (Setup Environment), base model download, and running detection are all covered there.
You’ve reached the end of this manual.Explore the library →