DetectionChinese & English

LocateAnything

LocateAnything is an object-localization (visual grounding) plugin in Dragonfly's Prototype Apps menu. You simply type what you are looking for (e.g. "pores", "cracks", "cells") and the plugin locates all matching object

Updated 2026-07-07User manual

LocateAnything(文字提示目标定位)插件用户手册

LocateAnything - User Manual

Dragonfly Prototype Apps · LocateAnything...

版本 Version 1.0 · 2026-07-04


第一部分 中文手册

目录

1. 简介

2. 适用场景

3. 安装与启用

4. 运行环境与首次配置

5. 界面说明

6. 使用步骤

7. 参数说明

8. 输出结果

9. 常见问题与故障排除

10. 注意事项与已知限制

11. 参考资料

1. 简介

LocateAnything 是 Dragonfly Prototype Apps 菜单中的目标定位(视觉定位 / grounding)插件。您只需输入想查找的目标名称(例如“孔隙”“裂纹”“细胞”),插件即可在图像的某一张二维切片上自动定位所有匹配对象。检测结果以 MultiROI 的形式返回 Dragonfly——每个被检测到的对象对应一个带标签的矩形区域——并尽力附加文字注释标签;同时在任务文件夹中保存一张画有检测框的叠加图(overlay.png),便于快速目视核查。

底层引擎为 NVIDIA LocateAnything-3B 视觉-语言定位大模型,通过 Hugging Face transformers 库(固定版本 4.57.1)加载,配合 CUDA 版 PyTorch 在本机 NVIDIA 显卡上运行推理。推理完全在插件自带的独立 Python 虚拟环境(venv)中以子进程方式进行,与 Dragonfly 通过 JSON 文件交换数据,绝不向 Dragonfly 自身的 Python 环境安装任何组件。

请注意:LocateAnything 是“定位”模型而不是“分割”模型。它返回的是候选区域(矩形框)或点,而不是逐像素的掩膜。得到的 MultiROI 适合作为后续精细分割、测量或标注的起点。

许可证要点:插件代码随 Prototype Apps 完整安装包分发;模型权重 LocateAnything-3B 由 NVIDIA 以非商业许可证发布在 Hugging Face(nvidia/LocateAnything-3B),下载不需要注册账号或访问令牌,但请自行阅读并遵守其许可条款(尤其在商业场景下)。所依赖的 transformers、PyTorch 等库遵循各自的开源许可证。

该模型基于自然照片(场景、文档、界面截图等)训练。将其用于灰度 CT / 显微图像属于实验性质,结果可能偏弱甚至为空——请始终结合 overlay.png 叠加图核查检测质量。

2. 适用场景

LocateAnything 适合在不训练任何模型的前提下快速定位候选特征:只要能用一个名词描述目标,就可以立即得到候选区域。典型场景包括:

  • 材料科学与工业 CT:在切片中查找孔隙(pores)、裂纹(cracks)、夹杂物(inclusions)等缺陷候选。
  • 生命科学显微图像:定位细胞、染色结构等目标;支持把最多三个灰度通道合成一张伪彩色图像输入模型(例如多通道荧光数据)。
  • 精细分割前的预筛选:先用文字提示圈出感兴趣区域,再对这些区域做精细分割、测量或人工标注。
  • 快速勘察:对陌生数据集,输入几个候选名词即可大致了解目标分布。

由于模型的训练数据是自然照片,它对灰度科学图像的定位能力是本插件重点评估的对象:结果弱或为空本身也是有价值的信息,并不一定意味着操作有误。

3. 安装与启用

本插件通过 Prototype Apps 完整安装包(Full Package) 安装:

1. 把完整安装包 zip 解压到任意位置(建议较短的路径,如 C:\PL\),双击 `Install_FullPackage.bat`。

2. 在弹出的安装对话框中勾选 LocateAnything(所有插件默认不勾选,需手动勾上)。核心安装模式选 Fresh 或 Compatible 均可——它只影响 Prototype Labs 核心的 blocks 与 recipes,不影响任何插件的环境与设置。

3. 点击 Install,等待控制台完成。

4. 完全退出并重启 Dragonfly。

5. 重启后,菜单栏出现 Prototype Apps ▸ LocateAnything...(位于 Detection 检测分组),点击即可打开一个可停靠的浮动面板。

以后想启用/停用:在 Dragonfly 中打开 Developer ▸ Prototype Labs... ▸ Menu Item Manager,在底部“Prototype Apps (Full Package)”列表中勾选或取消 LocateAnything,重启 Dragonfly 生效。停用从不删除插件已搭好的 venv 环境,重新启用立即可用。也可以随时重跑安装器(zip 删除后仍可运行 %LOCALAPPDATA%\DragonflyPrototypeLabs\FullPackage\installer\Install_FullPackage.bat),上次的勾选就是新的默认值。

升级:重跑安装器时保持“Update already-installed plugin code”复选框勾选(默认勾选),插件代码会刷新,同时保留已建好的 venv、已保存的配置和任务文件夹。卸载:双击 Uninstall_FullPackage.bat,它移除所有 Full Package 菜单项与插件,但保留各插件环境(结束时会列出路径,需要腾磁盘空间时可手动删除)。

Dragonfly 只在启动时扫描菜单——每次修改勾选(启用/停用/安装/卸载)后都需要重启一次 Dragonfly。

4. 运行环境与首次配置

运行本插件需要满足以下条件:

  • 一块 NVIDIA 显卡及较新驱动(建议 RTX 3080 / Ampere 架构或更新,显存 8 GB 以上)。
  • 网络连接:首次搭建环境需从 PyTorch 与 PyPI 源下载依赖;下载模型权重需访问 Hugging Face。推理本身在本机离线运行。
  • 磁盘空间:venv 依赖约数 GB,模型权重约 6 GB。
  • 不需要单独安装 Python:环境默认基于 Dragonfly 自带的 Python 构建。

4.1 Setup Environment 做了什么

首次使用时,在面板左侧点击 Setup Environment (build venv + install)(一次性操作,约 5–15 分钟、下载数 GB)。它依次执行:

1. 在插件代码目录下创建虚拟环境 venv(位于 %LOCALAPPDATA%\comet\<Dragonfly版本>\pythonUserExtensions\GenericMenuItems\LocateAnything\venv)。基础解释器默认为本 Dragonfly 自带的 python.exe,因此版本始终匹配;“Base Python (build)”字段可覆盖。

2. 在 venv 内升级 pip / setuptools / wheel。

3. 从“Torch CUDA wheel index”(默认 https://download.pytorch.org/whl/cu124)安装 CUDA 版 torch + torchvision。

4. 安装其余依赖:transformers==4.57.1、accelerate、safetensors、einops、sentencepiece、pillow、opencv-python-headless、timm、peft、lmdb、decord、numpy 等。

5. 自检:在 venv 中导入 torch 与 transformers,并在日志中以 SMOKE=... 行报告 CUDA 是否可用及显卡名称。

6. 成功后,“Inference venv python”字段自动填入 venv 的 python 路径并保存——以后不必重复此步骤。

4.2 下载模型权重(一次性)

权重不随插件分发,需自行下载一次(约 6 GB,无需登录或令牌)。在普通命令行窗口(不是 Dragonfly 内)中用 venv 自带的工具执行:

<venv>\Scripts\huggingface-cli download nvidia/LocateAnything-3B --local-dir C:\models\LocateAnything-3B

其中 <venv> 即上一步创建的虚拟环境目录(%LOCALAPPDATA%\comet\<Dragonfly版本>\pythonUserExtensions\GenericMenuItems\LocateAnything\venv)。下载完成后,把面板的 Model folder 指向直接包含 `config.json` 的文件夹(如 C:\models\LocateAnything-3B)。

请把权重放在 Dragonfly 目录之外(例如 C:\models\LocateAnything-3B):路径中含 “Dragonfly” 字样在配置文件读取异常时可能干扰 transformers 的模型类型识别;并且放在插件目录内的文件会随插件卸载被一并删除。

4.3 相关文件位置

  • 推理环境 venv:插件代码目录内的 venv\(见 4.1)。
  • 任务输出:默认 C:\LocateJobs(面板 Job root 可改),每次运行生成一个 loc_<日期_时间> 子文件夹。
  • 已保存设置:插件代码目录的 locateanything_config.json 以及 %LOCALAPPDATA%\LocateAnything\config.json(后者优先,重装插件也不丢失)。

4.4 失败时的替代方案

  • torch 安装失败:检查网络,并确认 CUDA wheel 源(默认 cu124)与显卡驱动匹配,必要时把索引地址改为其他 CUDA 版本(如 cu121)后重试。
  • venv 创建失败:把 “Base Python (build)” 指向本机任意 CPython 3.9+(例如 C:\Python312\python.exe,或命令形式 py -3.12)后重试。重跑 Setup 会自动识别“建了一半”的 venv(有 python 但 pip 不可用)并重建。
  • Windows 下模型无法加载:把 Run mode 切换为 wsl(实验性备选,需自行准备 WSL 发行版及 Linux 侧的 Python 环境,面向高级用户)。

5. 界面说明

面板为左右两栏布局,中间的分隔条可以拖动。左栏是全部配置项(带独立滚动条),从上到下分为 Environment & Model、Input Slice、Prompt、Output 四组;右栏是运行按钮、结果行和日志。面板顶部有一行蓝色说明文字,提醒模型的实验性质。

5.1 Environment & Model(环境与模型)

  • Model folder:本地 LocateAnything-3B 权重文件夹(必须直接包含 config.json)。右侧 “…” 按钮打开文件夹选择框;选定后立即保存,重启后仍然记住。
  • Inference venv python:推理 venv 的 python.exe 路径。由 Setup Environment 成功后自动填写,一般无需手改。
  • Base Python (build):仅在搭建环境时使用的基础解释器。留空 = 使用本 Dragonfly 自带的 python.exe(推荐);也可填某个 python.exe 的完整路径或命令(如 py -3.12)。
  • Run mode:windows(默认)或 wsl(实验性备选,见 4.4)。
  • Torch CUDA wheel index:安装 torch 使用的 pip 源,默认 https://download.pytorch.org/whl/cu124。
  • Job root:任务输出根目录,默认 C:\LocateJobs。
  • Setup Environment (build venv + install) 按钮:一次性搭建推理环境(见 4.1)。

5.2 Input Slice(输入切片——三个灰度通道合成 RGB)

  • Red channel + Refresh 按钮:红色通道是主通道——它决定几何信息与切片范围,检测结果也发布在它的网格上。点 Refresh 重新扫描当前工程中的全部 Channel。
  • 选定 Red 后,面板会按名称自动匹配 Green/Blue 同组通道(例如 “Red Ab_06” 自动填 “Green Ab_06”“Blue Ab_06”;“food_R” 匹配 “food_G”;“ch_red” 匹配 “ch_green”)。
  • Green channel / Blue channel:可手动更改;选 “(none)” 表示该颜色平面填 0。三个灰度通道合成一张伪彩色 RGB 图输入模型;普通灰度数据把三个下拉都选同一通道即可。
  • Orientation:切片方向,XY / XZ / YZ,默认 XY。
  • Slice index:切片序号。范围随 Red 通道尺寸与方向自动更新,首次自动定位到中间切片。
  • Intensity window:灰度窗下限 lo / 上限 hi。留空(显示 “auto”)= 自动确定;也可手动输入数值(须 hi > lo 才生效)。
  • Max image size (px):导出图像的最长边上限,范围 256–4096、步长 128,默认 1024。超过上限的切片会等比缩小后再送入模型。

5.3 Prompt(提示词)

  • Task:任务类型,detect / ground / ground_multi / point / detect_text,默认 detect。任务决定发送给模型的提示模板(例如 detect 对应 “Locate all instances matching: <提示词>”);point 返回点而非框;detect_text 检测图中的文字,不需要填提示词。
  • Prompt:要查找的对象/类别。应填名词(如 pores、cracks、cells),不要写成命令句——模板本身已含 “Locate all ...” 动词,写 “locate all pores” 会变成动词重复的句子。代码默认值为 pores;面板会记住您上次保存的值。
  • Max new tokens:模型生成的最大 token 数,范围 64–8192、步长 64,默认 1024。检测对象非常多时可适当调大。

5.4 Output(输出选项)

  • Create MultiROI (one labeled ROI per detection):默认勾选。把每个检测导入为 MultiROI 中的一个带标签区域。
  • Add annotation labels (best-effort):默认勾选。尽力为每个检测添加文字注释标签;若当前 Dragonfly 版本不提供相应注释接口则自动跳过(日志会说明),不影响 MultiROI。
  • Save overlay.png (boxes drawn on the input):默认勾选。在任务文件夹保存画有检测框的叠加图。

5.5 右栏:运行与日志

  • Run LocateAnything(蓝色按钮):开始一次运行。运行前所有当前设置都会被保存。
  • Open Output Folder:在资源管理器中打开最近一次任务文件夹(若尚未运行则打开 Job root)。
  • 结果行:完成后显示 “Done. N detection(s)...”,包括 MultiROI 标签数、注释数和叠加图文件名;失败时显示错误原因。
  • Log:完整过程日志(导出切片、venv 搭建输出、推理进度、警告与错误)。

6. 使用步骤

前提:已按第 3 章安装启用插件,并按第 4 章完成一次 Setup Environment 与权重下载。

1. 在 Dragonfly 中加载图像数据(Channel)。

2. 打开 Prototype Apps ▸ LocateAnything...。

3. 确认 Model folder(指向含 config.json 的权重文件夹)与 Inference venv python(由 Setup 自动填写)均已就绪。

4. 点 Refresh,在 Red channel 中选择要分析的通道。普通灰度数据:Green/Blue 会自动跟随或可手动选同一通道 / “(none)”;多通道彩色数据:分别选择三个通道合成伪彩色图。

5. 选择 Orientation(XY/XZ/YZ)和 Slice index(默认已定位到中间切片)。需要时手动填写 Intensity window,否则保持自动。

6. 选择 Task(常用 detect),在 Prompt 中输入名词类提示词,例如 pores。

7. 按需勾选三个输出选项,点击 Run LocateAnything。

8. 观察右侧日志:导出切片 → 启动推理子进程(首次加载模型较慢)→ 返回检测数量。

9. 完成后:结果行显示 “Done. N detection(s)...”。在 Dragonfly 对象列表中找到新的 MultiROI(名为 LocateAnything: <提示词>)并叠加到视图查看;点 Open Output Folder 打开任务文件夹核对 overlay.png。

10. 若检测为 0 或只有覆盖全图的大框:换一个更简短的名词提示词、试 point 任务、换一张结构更清晰的切片、或调整灰度窗后重试(日志中也会打印同样的建议)。

7. 参数说明

参数

默认值

说明

Model folder

(空)

本地 LocateAnything-3B 权重文件夹,必须直接包含 config.json;选定后立即保存

Inference venv python

(空,由 Setup 自动填写)

推理 venv 的 python.exe 路径;未填时无法运行

Base Python (build)

(空 = Dragonfly 自带 python)

仅搭建环境时使用;可填 python.exe 路径或命令(如 py -3.12)

Run mode

windows

windows / wsl;wsl 为实验性备选(高级用户)

Torch CUDA wheel index

https://download.pytorch.org/whl/cu124

安装 torch 的 pip 源;须与显卡驱动的 CUDA 版本匹配

Job root

C:\LocateJobs

任务输出根目录;每次运行生成 loc_<日期_时间> 子文件夹

Red channel

(第一次需手动选择)

主通道:决定几何、切片范围与结果发布位置

Green / Blue channel

按名称自动匹配

可手动改;(none) = 该颜色平面为 0

Orientation

XY

切片方向:XY / XZ / YZ

Slice index

中间切片

0 到 N-1,范围随通道尺寸与方向自动更新

Intensity window (lo/hi)

auto(留空)

灰度窗;两项都填且 hi > lo 时才生效

Max image size (px)

1024

导出图像最长边上限;范围 256–4096,步长 128

Task

detect

detect / ground / ground_multi / point / detect_text;detect_text 无需提示词

Prompt

pores

要查找的对象/类别,应为名词而非命令句;记住上次保存值

Max new tokens

1024

生成上限;范围 64–8192,步长 64

Create MultiROI

勾选

每个检测导入为一个带标签 ROI

Add annotation labels

勾选

尽力添加文字注释;接口不可用时自动跳过

Save overlay.png

勾选

在任务文件夹保存画框叠加图

以上设置在每次运行、每次环境搭建、选择模型文件夹以及关闭面板时都会自动保存,重启 Dragonfly 后自动恢复。

8. 输出结果

一次成功的运行会在 Dragonfly 内和任务文件夹中分别产生以下内容:

  • MultiROI:名为 LocateAnything: <提示词前 40 字符>,发布在 Red 通道的几何网格上;每个检测对象对应一个标签,矩形区域绘制在所选切片上。在对象列表中勾选即可叠加显示,并可直接用于后续分割、测量或统计。
  • 注释标签(Annotation):每个检测一个,标题为 <MultiROI 名> (det);仅在当前 Dragonfly 版本提供注释接口时创建(否则日志中说明并跳过)。
  • 任务文件夹 C:\LocateJobs\loc_<日期_时间>\:input.npy(导出的切片数据)、export.json(几何与逆变换记录,便于追溯)、config.json(发给推理器的完整配置)、results.json(检测明细:box/point 类型、标签、归一化坐标与像素坐标、原始模型回答、耗时与显存占用)、status.json(运行状态)、overlay.png(画框叠加图)。

查看方式:点 Open Output Folder 打开任务文件夹,双击 overlay.png 即可快速核查检测框位置是否合理;MultiROI 在 Dragonfly 中按常规 ROI 操作查看与编辑。

9. 常见问题与故障排除

问:菜单里没有 “LocateAnything...”?

答:安装完整安装包时未勾选该插件(插件默认不勾选),或修改勾选后没有重启 Dragonfly。到 Developer ▸ Prototype Labs... ▸ Menu Item Manager 勾选后重启即可。

问:Setup Environment 失败(torch 安装报错)?

答:先检查网络;再确认 “Torch CUDA wheel index”(默认 cu124)与本机显卡驱动匹配,必要时改用其他 CUDA 版本的索引地址(如 cu121)重试。若是 venv 创建本身失败,把 “Base Python (build)” 指向另一个 CPython 3.9+ 后重跑。重跑 Setup 会自动检测“有 python 但 pip 不可用”的半成品 venv 并重建。

问:点 Run 提示 “Inference venv not set. Click 'Setup Environment' first.”?

答:还没有搭建推理环境。点一次 Setup Environment,成功后 “Inference venv python” 会自动填好。

问:报错 “Model folder ... has no config.json”?

答:Model folder 指向的层级不对或权重尚未下载完整。它必须指向直接包含 `config.json` 的那个文件夹(例如 C:\models\LocateAnything-3B),而不是它的上层目录。

问:运行失败并提示 CUDA 不可用或显存不足(no_cuda / cuda_oom)?

答:确认本机有 NVIDIA 显卡且驱动较新;关闭其他占用显存的程序;把 Max image size 调小后重试。显存建议 8 GB 以上。

问:检测数为 0,或只返回一两个覆盖整幅图的大框?

答:说明模型没有真正定位到目标(日志会有对应提示)。改用简短名词作提示词(如 pores,而不是 “locate all pores” 这类命令句);试试 point 任务;换一张结构更清晰的切片;或手动调整灰度窗。灰度 CT/显微数据超出了模型的训练分布,结果偏弱属正常现象,不一定是故障。

问:每次 Run 前面等待很久?

答:每次运行都会启动独立的推理子进程并把约 6 GB 的模型加载进显存,加载时间占比较高属于正常现象;日志会实时显示进度。

10. 注意事项与已知限制

  • 模型训练于自然图像;用于灰度 CT / 显微数据为实验性,请以 overlay.png 为准判断质量。空结果也是有效信息,不一定是 bug。
  • 定位不是分割:输出为矩形框或点,不是逐像素掩膜;MultiROI 中每个标签是绘制在所选切片上的矩形区域。
  • 一次只处理一张 2D 切片;暂不支持 3D 全体积逐片扫描。
  • 检测结果的置信度字段可能为空(模型不总是输出置信度)。
  • 模型权重为 NVIDIA 非商业许可证;商业用途请自行确认授权。
  • 只有环境搭建与权重下载两个阶段需要联网;推理在本机离线运行。
  • wsl 运行模式是面向高级用户的实验性备选,需要自行准备 WSL 发行版及 Linux 侧环境。
  • 权重文件夹不要放在路径含 “Dragonfly” 字样的位置,也不要放进插件目录内(见 4.2 的说明)。
  • 在 Menu Item Manager 中停用插件从不删除其环境;Full Package 卸载器也保留插件环境(venv、下载内容),结束时列出路径供手动清理。

11. 参考资料

  • 模型权重与模型卡:Hugging Face 上的 nvidia/LocateAnything-3B(含许可证与推荐用法)。
  • NVIDIA “LocateAnything” 项目页(模型的官方介绍)。
  • 推理框架:Hugging Face Transformers(本插件固定使用 4.57.1)与 PyTorch(CUDA 版,默认 cu124 wheel 源)。
  • Prototype Apps 完整安装包自述文件:安装、启用/停用(Menu Item Manager)、卸载与“路径过长”问题的完整说明。


Part II English Manual

Contents

1. Overview

2. Typical Use Cases

3. Installation & Enabling

4. Runtime Environment & First-Time Setup

5. User Interface

6. Step-by-Step Usage

7. Parameter Reference

8. Outputs

9. FAQ & Troubleshooting

10. Notes & Known Limitations

11. References

1. Overview

LocateAnything is an object-localization (visual grounding) plugin in Dragonfly's Prototype Apps menu. You simply type what you are looking for (e.g. "pores", "cracks", "cells") and the plugin locates all matching objects on a chosen 2D slice of your image. The detections come back into Dragonfly as a MultiROI — one labeled rectangular region per detected object — plus best-effort text annotation labels, and an overlay.png with the detection boxes drawn on the input is saved in the job folder for a quick visual check.

The underlying engine is the NVIDIA LocateAnything-3B vision-language grounding model, loaded through the Hugging Face transformers library (pinned to 4.57.1) with CUDA PyTorch running on your local NVIDIA GPU. Inference runs entirely in the plugin's own dedicated Python virtual environment (venv) as a subprocess, exchanging data with Dragonfly via JSON files — nothing is ever installed into Dragonfly's own Python.

Note that LocateAnything is a *localization* model, not a segmentation model: it returns candidate regions (boxes) or points, not per-pixel masks. The resulting MultiROI is intended as a starting point for further segmentation, measurement, or annotation.

Licensing: the plugin code ships with the Prototype Apps Full Package. The LocateAnything-3B weights are published by NVIDIA on Hugging Face (nvidia/LocateAnything-3B) under an NVIDIA non-commercial license; the download requires no account or token, but please read and respect the license terms, especially for commercial use. Dependencies (transformers, PyTorch, ...) carry their own open-source licenses.

The model was trained on natural photographs (scenes, documents, GUIs, ...). Grounding on grayscale CT/microscopy data is experimental — results may be weak or empty. Always check the overlay.png.

2. Typical Use Cases

LocateAnything is useful whenever you want to quickly locate candidate features without training any model: if you can name the target with a noun, you can get candidate regions immediately. Typical scenarios:

  • Materials science and industrial CT: finding pores, cracks, or inclusions in slices as defect candidates.
  • Life-science microscopy: spotting cells or stained structures; up to three grayscale channels can be composited into one false-color image for the model (e.g. multi-channel fluorescence data).
  • Pre-selection before detailed segmentation: outline regions of interest from a text prompt, then segment, measure, or annotate those regions precisely.
  • Quick surveys: on an unfamiliar dataset, try a few candidate nouns to get a rough idea of where targets are.

Because the training data are natural photographs, the model's performance on grayscale scientific images is exactly what this plugin lets you evaluate: weak or empty results are themselves useful information, not necessarily a usage error.

3. Installation & Enabling

The plugin installs via the Prototype Apps Full Package:

1. Unzip the Full Package anywhere (prefer a short path such as C:\PL\) and double-click `Install_FullPackage.bat`.

2. In the installer dialog, tick LocateAnything (all plugins are unticked by default). Either core mode (Fresh or Compatible) is fine — it only affects the Prototype Labs core blocks & recipes, never any plugin's environment or settings.

3. Click Install and wait for the console to finish.

4. Quit Dragonfly completely and restart it.

5. After the restart, Prototype Apps ▸ LocateAnything... appears in the menu bar (in the Detection group). Clicking it opens a dockable floating panel.

To enable/disable later: open Developer ▸ Prototype Labs... ▸ Menu Item Manager inside Dragonfly — the "Prototype Apps (Full Package)" list at the bottom has a checkbox per app; tick = deploy, untick = remove the menu entry, then restart Dragonfly. Disabling never deletes the plugin's built venv — re-enabling is instant. You can also re-run the installer anytime (even after deleting the zip: %LOCALAPPDATA%\DragonflyPrototypeLabs\FullPackage\installer\Install_FullPackage.bat); your previous choices are the new defaults.

Updating: when re-running the installer, keep "Update already-installed plugin code" ticked (the default) — the plugin's code is refreshed while its venv, saved settings, and job folders are preserved. Uninstalling: double-click Uninstall_FullPackage.bat; it removes all Full-Package menu items and plugins but keeps every plugin environment (the paths are listed at the end so you can delete them manually to reclaim disk space).

Dragonfly scans its menus only at startup — every enable/disable/install/uninstall change needs one Dragonfly restart.

4. Runtime Environment & First-Time Setup

Requirements:

  • An NVIDIA GPU with a recent driver (RTX 3080 / Ampere or newer recommended; 8 GB+ VRAM recommended).
  • Internet access: the one-time environment setup downloads packages from the PyTorch and PyPI indexes, and the model weights are downloaded from Hugging Face. Inference itself runs locally and offline.
  • Disk space: a few GB for the venv, plus about 6 GB for the model weights.
  • No separate Python install needed: the environment is built from Dragonfly's own bundled Python by default.

4.1 What "Setup Environment" does

On first use, click Setup Environment (build venv + install) on the left side of the panel (one-time, roughly 5-15 minutes and a few GB of downloads). It performs, in order:

1. Creates the virtual environment venv inside the plugin's code folder (%LOCALAPPDATA%\comet\<DragonflyVersion>\pythonUserExtensions\GenericMenuItems\LocateAnything\venv). The base interpreter defaults to this Dragonfly's own python.exe, so versions always match; the "Base Python (build)" field overrides this.

2. Upgrades pip / setuptools / wheel inside the venv.

3. Installs CUDA torch + torchvision from the "Torch CUDA wheel index" (default https://download.pytorch.org/whl/cu124).

4. Installs the remaining stack: transformers==4.57.1, accelerate, safetensors, einops, sentencepiece, pillow, opencv-python-headless, timm, peft, lmdb, decord, numpy, etc.

5. Runs a smoke test: imports torch and transformers in the venv and reports CUDA availability and the GPU name in the log (SMOKE=... line).

6. On success, the "Inference venv python" field is filled in automatically and saved — you never need to repeat this step.

4.2 Downloading the model weights (one-time)

The weights are not shipped with the plugin; download them once (about 6 GB, no login or token needed). In a normal command prompt (not inside Dragonfly), use the venv's own tool:

<venv>\Scripts\huggingface-cli download nvidia/LocateAnything-3B --local-dir C:\models\LocateAnything-3B

Here <venv> is the environment created in the previous step (%LOCALAPPDATA%\comet\<DragonflyVersion>\pythonUserExtensions\GenericMenuItems\LocateAnything\venv). When the download finishes, point the panel's Model folder at the folder that directly contains `config.json` (e.g. C:\models\LocateAnything-3B).

Keep the weights outside the Dragonfly folders (e.g. C:\models\LocateAnything-3B): a path containing the word "Dragonfly" can confuse transformers' model-type guesser if the config file is ever unreadable, and files stored inside the plugin folder would be deleted together with the plugin on uninstall.

4.3 Where things live

  • Inference venv: venv\ inside the plugin's code folder (see 4.1).
  • Job output: C:\LocateJobs by default (configurable via Job root); each run creates a loc_<date_time> subfolder.
  • Saved settings: locateanything_config.json next to the plugin code, plus %LOCALAPPDATA%\LocateAnything\config.json (the latter wins and survives plugin re-installs).

4.4 If something fails

  • torch install fails: check your internet connection, and make sure the CUDA wheel index (default cu124) matches your GPU driver; switch the index URL to another CUDA build (e.g. cu121) and retry.
  • venv creation fails: point "Base Python (build)" at any CPython 3.9+ on your machine (e.g. C:\Python312\python.exe, or a command like py -3.12) and retry. Re-running Setup automatically detects a half-built venv (python present but pip broken) and rebuilds it.
  • The model will not load on Windows: switch Run mode to wsl (an experimental fallback for advanced users; you must provide a WSL distro with a Linux-side Python environment yourself).

5. User Interface

The panel has a two-column layout with a draggable splitter. The left column holds all configuration (with its own scrollbar), grouped top-to-bottom into Environment & Model, Input Slice, Prompt, and Output. The right column holds the run buttons, the result line, and the log. A blue note at the top reminds you of the model's experimental nature on scientific images.

5.1 Environment & Model

  • Model folder: the local LocateAnything-3B weights folder (must directly contain config.json). The "…" button opens a folder picker; the choice is saved immediately and remembered across restarts.
  • Inference venv python: path to the inference venv's python.exe. Filled in automatically by a successful Setup Environment; normally never edited by hand.
  • Base Python (build): the base interpreter used only for *building* the venv. Blank = this Dragonfly's own python.exe (recommended); or a full path to a python.exe, or a command like py -3.12.
  • Run mode: windows (default) or wsl (experimental fallback, see 4.4).
  • Torch CUDA wheel index: pip index used to install torch, default https://download.pytorch.org/whl/cu124.
  • Job root: root folder for job output, default C:\LocateJobs.
  • Setup Environment (build venv + install) button: one-time environment build (see 4.1).

5.2 Input Slice (RGB composite from 3 grayscale channels)

  • Red channel + Refresh button: Red is the primary channel — it drives the geometry and slice range, and the results are published on its grid. Refresh re-scans all Channels in the current project.
  • After picking Red, the panel auto-fills Green/Blue with same-named sibling channels (e.g. "Red Ab_06" auto-fills "Green Ab_06" / "Blue Ab_06"; "food_R" matches "food_G"; "ch_red" matches "ch_green").
  • Green channel / Blue channel: editable; choose "(none)" to fill that color plane with zeros. The three grayscale channels are composited into one false-color RGB image for the model; for plain grayscale data simply select the same channel in all three (or leave the auto-fill).
  • Orientation: slice plane, XY / XZ / YZ, default XY.
  • Slice index: slice number. The range follows the Red channel's shape and the orientation; on first populate it jumps to the middle slice.
  • Intensity window: gray-window lo / hi. Leave blank ("auto") for automatic windowing; manual values only take effect when both are filled and hi > lo.
  • Max image size (px): cap on the exported image's longest side, range 256-4096, step 128, default 1024. Larger slices are downscaled proportionally before being sent to the model.

5.3 Prompt

  • Task: detect / ground / ground_multi / point / detect_text, default detect. The task selects the prompt template sent to the model (e.g. detect maps to "Locate all instances matching: <prompt>"); point returns points instead of boxes; detect_text detects text in the image and needs no prompt.
  • Prompt: the object/category to find. Use a noun (e.g. pores, cracks, cells), not a command — the templates already contain the "Locate all ..." verb, so typing "locate all pores" produces a doubled-verb sentence. The code default is pores; the panel remembers your last saved value.
  • Max new tokens: generation cap for the model, range 64-8192, step 64, default 1024. Increase it when many objects are expected.

5.4 Output

  • Create MultiROI (one labeled ROI per detection): checked by default. Imports each detection as one labeled region in a MultiROI.
  • Add annotation labels (best-effort): checked by default. Adds a text annotation per detection when the running Dragonfly version provides the annotation interface; otherwise it is skipped with a log message, without affecting the MultiROI.
  • Save overlay.png (boxes drawn on the input): checked by default. Saves the overlay image into the job folder.

5.5 Right column: run & log

  • Run LocateAnything (blue button): starts a run. All current settings are saved before the run.
  • Open Output Folder: opens the most recent job folder in Explorer (or the Job root if nothing has run yet).
  • Result line: after completion shows "Done. N detection(s)..." with the MultiROI label count, annotation count, and overlay file name; on failure it shows the error.
  • Log: the full progress log (slice export, venv build output, inference progress, warnings and errors).

6. Step-by-Step Usage

Prerequisites: the plugin is installed and enabled (Chapter 3), and Setup Environment plus the one-time weights download are done (Chapter 4).

1. Load your image data (a Channel) in Dragonfly.

2. Open Prototype Apps ▸ LocateAnything....

3. Confirm that Model folder (pointing at the folder containing config.json) and Inference venv python (auto-filled by Setup) are both set.

4. Click Refresh and pick the Red channel to analyze. Plain grayscale data: Green/Blue follow automatically, or select the same channel / "(none)" manually; multi-channel color data: pick the three channels to composite a false-color image.

5. Choose the Orientation (XY/XZ/YZ) and the Slice index (it defaults to the middle slice). Optionally set the Intensity window, or leave it on auto.

6. Pick a Task (usually detect) and type a noun-style Prompt, e.g. pores.

7. Tick the output options you want, then click Run LocateAnything.

8. Watch the log on the right: slice export, then the inference subprocess starts (loading the model takes a while, especially the first time), then the number of detections is reported.

9. When it finishes: the result line shows "Done. N detection(s)...". Find the new MultiROI (named LocateAnything: <prompt>) in Dragonfly's object list and overlay it on the view; click Open Output Folder to inspect overlay.png.

10. If you get 0 detections or only whole-image boxes: try a shorter noun prompt, the point task, a slice with clearer structures, or a manual intensity window, then run again (the log prints the same tips).

7. Parameter Reference

Parameter

Default

Description

Model folder

(empty)

Local LocateAnything-3B weights folder; must directly contain config.json; saved as soon as picked

Inference venv python

(empty; filled by Setup)

Path to the inference venv's python.exe; runs are refused while unset

Base Python (build)

(empty = Dragonfly's own python)

Used only when building the venv; a python.exe path or a command such as py -3.12

Run mode

windows

windows / wsl; wsl is an experimental fallback for advanced users

Torch CUDA wheel index

https://download.pytorch.org/whl/cu124

pip index for the torch install; must match your GPU driver's CUDA level

Job root

C:\LocateJobs

Root output folder; each run creates a loc_<date_time> subfolder

Red channel

(pick on first use)

Primary channel: drives geometry, slice range, and where results are published

Green / Blue channel

auto-matched by name

Editable; (none) = that color plane is zero

Orientation

XY

Slice plane: XY / XZ / YZ

Slice index

middle slice

0 to N-1; range follows the channel shape and orientation

Intensity window (lo/hi)

auto (blank)

Gray window; effective only when both values are set and hi > lo

Max image size (px)

1024

Cap on the exported image's longest side; range 256-4096, step 128

Task

detect

detect / ground / ground_multi / point / detect_text; detect_text needs no prompt

Prompt

pores

Object/category to find; a noun, not a command; last saved value is remembered

Max new tokens

1024

Generation cap; range 64-8192, step 64

Create MultiROI

checked

Import each detection as one labeled ROI

Add annotation labels

checked

Best-effort text annotations; skipped if the interface is unavailable

Save overlay.png

checked

Save the boxed overlay image into the job folder

All settings are saved automatically on every run, on every environment setup, when the model folder is picked, and when the panel closes — they are restored after a Dragonfly restart.

8. Outputs

A successful run produces the following, inside Dragonfly and in the job folder:

  • MultiROI: named LocateAnything: <first 40 chars of the prompt>, published on the Red channel's grid; one label per detected object, drawn as a rectangular region on the chosen slice. Tick it in the object list to overlay it, and use it directly for further segmentation, measurement, or statistics.
  • Annotation labels: one per detection, titled <MultiROI name> (det); created only when the running Dragonfly version provides the annotation interface (otherwise skipped with a log message).
  • Job folder C:\LocateJobs\loc_<date_time>\: input.npy (the exported slice), export.json (geometry / inverse-transform record for traceability), config.json (the full config sent to the runner), results.json (detection details: box/point type, label, normalized + pixel coordinates, the raw model answer, timing and VRAM usage), status.json (run state), and overlay.png (the boxed overlay image).

To view the results: click Open Output Folder and double-click overlay.png for a quick sanity check of the box positions; the MultiROI is viewed and edited in Dragonfly like any other ROI object.

9. FAQ & Troubleshooting

Q: "LocateAnything..." does not appear in the menu?

A: The plugin was not ticked during the Full Package install (plugins are unticked by default), or Dragonfly was not restarted after a change. Tick it in Developer ▸ Prototype Labs... ▸ Menu Item Manager and restart.

Q: Setup Environment fails (torch install error)?

A: Check your internet connection first; then make sure the "Torch CUDA wheel index" (default cu124) matches your GPU driver, switching to another CUDA build (e.g. cu121) if needed. If the venv creation itself fails, point "Base Python (build)" at another CPython 3.9+ and retry. Re-running Setup detects a half-built venv (python present, pip broken) and rebuilds it automatically.

Q: Run says "Inference venv not set. Click 'Setup Environment' first."?

A: The inference environment has not been built yet. Click Setup Environment once; on success the "Inference venv python" field is filled automatically.

Q: Error "Model folder ... has no config.json"?

A: The Model folder points at the wrong level, or the weights download is incomplete. It must point at the folder that directly contains `config.json` (e.g. C:\models\LocateAnything-3B), not its parent.

Q: The run fails with CUDA unavailable or out of memory (no_cuda / cuda_oom)?

A: Verify you have an NVIDIA GPU with a recent driver; close other VRAM-hungry applications; reduce Max image size and retry. 8 GB+ VRAM is recommended.

Q: Zero detections, or only one or two boxes covering the whole image?

A: The model did not actually localize anything (the log prints a corresponding note). Use a short noun as the prompt (e.g. pores, not a command like "locate all pores"); try the point task; pick a slice with clearer structures; or set a manual intensity window. Grayscale CT/microscopy is outside the model's training distribution, so weak results are expected behavior, not necessarily a fault.

Q: Every run has a long wait before results appear?

A: Each run launches a fresh inference subprocess and loads the ~6 GB model into GPU memory, so the loading phase dominates — this is normal; the log shows live progress.

10. Notes & Known Limitations

  • The model was trained on natural images; use on grayscale CT/microscopy is experimental. Judge quality from overlay.png; empty results are valid data, not necessarily a bug.
  • Localization is not segmentation: outputs are boxes or points, not per-pixel masks; each MultiROI label is a rectangle painted on the chosen slice.
  • Only one 2D slice is processed per run; there is no 3D whole-volume sweep yet.
  • The confidence field of a detection may be empty (the model does not always output one).
  • The model weights carry an NVIDIA non-commercial license; verify the licensing yourself before commercial use.
  • Internet is needed only for the environment build and the weights download; inference runs locally and offline.
  • The wsl run mode is an experimental fallback for advanced users and requires a self-prepared WSL distro with a Linux-side environment.
  • Do not place the weights folder under a path containing the word "Dragonfly", and do not place it inside the plugin folder (see 4.2).
  • Disabling the plugin in the Menu Item Manager never deletes its environment; the Full Package uninstaller also keeps plugin environments (venv, downloads) and lists their paths for manual cleanup.

11. References

  • Model weights and model card: nvidia/LocateAnything-3B on Hugging Face (includes the license and recommended usage).
  • The NVIDIA "LocateAnything" project page (the model's official description).
  • Inference stack: Hugging Face Transformers (pinned to 4.57.1 by this plugin) and PyTorch (CUDA build, default cu124 wheel index).
  • The Prototype Apps Full Package README: complete instructions for installing, enabling/disabling (Menu Item Manager), uninstalling, and the "path too long" issue.
You’ve reached the end of this manual.Explore the library →