Cloud GPUChinese & English

Train/Infer Cellpose on Cloud Server

This plugin runs both Cellpose training and inference on a cloud server's free GPU, with the same Train (left) | Infer (right) two-column layout as the built-in Cellpose plugin plus a Cloud server dropdown. Use it when t

Updated 2026-07-19User manual

在云端服务器上训练/推理 Cellpose 模型

Train/Infer Cellpose on Cloud Server - User Manual

Dragonfly Prototype Apps · Train/Infer Cellpose on Cloud Server...

版本 Version 1.4 · 2026-07-19


第一部分 中文手册

目录

1. 简介

2. 云服务器与模式

3. 账号与凭据设置(自动模式必读)

4. 远程 Dragonfly 服务器(第 5 个选项)

5. 安装

6. 使用:训练(左栏)

7. 使用:推理(右栏)

8. 常见问题排查

9. 环境需求与注意事项

1. 简介

本插件把 Cellpose 模型的训练与推理放到云端服务器的免费 GPU 上完成,界面采用与内置 Cellpose 插件一致的训练(左)| 推理(右)双栏布局,并在顶部提供一个 Cloud server(云服务器)下拉框。适用于本机没有独立显卡、或显存太小跑不动 Cellpose-SAM 的情况——本机只负责读取数据与发布结果,真正耗费显卡的计算在云端完成。

为什么需要在浏览器点一下:免费版云 notebook 没有公开的"无头运行"接口,GPU 计算必须在浏览器标签页里点 Run all。插件会自动生成好 notebook,并把其余环节(读数据、打包、传输、发布)都自动化。

2. 云服务器与模式

顶部下拉框选择五个平台之一,并选择 Auto(自动) 或 Manual(手动) 模式:

云服务器 / Host

自动 Auto

手动 Manual

凭据 / Credentials

Google Colab

✅

✅

OAuth client_secret.json 或 service_account.json

Kaggle

✅

✅

kaggle.json 文件,或 KGAT_ 令牌 + 用户名

百度飞桨 AI Studio

—

✅

无(手动)

阿里 ModelScope 魔搭

—

✅

无(手动)

Remote Dragonfly Server (via Prototype Apps)

✅

—

服务器 IP + 端口(+ 令牌 token)

  • 手动模式(前四家云 notebook 支持):插件把 input.zip 和对应平台的 notebook 导出到本地;你在浏览器里上传运行、下载 output.zip,再用"载入模型/载入结果"按钮导回。无需任何账号 API 或环境配置。
  • 自动模式(Colab、Kaggle 与 Remote Dragonfly Server):插件自动上传数据、执行计算、并下载结果。Colab/Kaggle 需要一次性配置账号凭据(见下一节);Remote Dragonfly Server 只需填服务器 IP/端口/令牌(见"远程 Dragonfly 服务器"一节)。
  • Remote Dragonfly Server(远程 Dragonfly 服务器):连接到另一台(或本机)运行着 Prototype Labs TCP 服务器的 Dragonfly,把训练/推理放到那台机器的 GPU 上跑。它用服务器自带的 Dragonfly Python(已内置 cellpose+torch),无需 notebook、无需 venv、无需 pip 安装,也不需要 Google/Kaggle 账号或翻墙。仅支持自动模式。

3. 账号与凭据设置(自动模式必读)

手动模式无需任何凭据,可跳过本节。 自动模式(Colab / Kaggle)需要按下面一次性设置好账号、API 令牌与权限。

Google Colab(自动)——OAuth 设置

需要一个 Google 账号。用 OAuth 客户端凭据让插件访问你自己的 Google 云端硬盘作为中转。

1. 打开 Google Cloud Console(console.cloud.google.com),新建或选择一个项目。

2. 在 API 和服务 ▸ 库 中启用 Google Drive API。

3. 配置 OAuth 同意屏幕(新版在 Google Auth Platform ▸ Audience):User type 选 External,填写应用名称与支持邮箱。

4. 在 Test users(测试用户) 中 添加你自己的 Gmail —— 这一步必须做,否则登录会被 Google 拦下并报 Access blocked … / Error 403: access_denied。

5. 在 凭据 ▸ 创建凭据 ▸ OAuth 客户端 ID 中,应用类型选 桌面应用 (Desktop app),创建后 下载 client_secret.json。

6. 在插件的 Google credentials 字段点 … 选择该 client_secret.json(下方会显示 detected: Google OAuth client)。

7. 首次点 Train/Infer on Google Colab 会弹出浏览器登录:用你(已加为 test user)的账号登录 → 出现 Google hasn't verified this app → 点 Advanced ▸ Go to <应用名> (unsafe) → Allow 授予 Google Drive 权限。令牌会被缓存,之后不再需要登录。

Share with (email) 字段:OAuth 模式用不到,请留空(它只用于服务账号模式)。也可以用服务账号 service_account.json,但个人 Gmail 不推荐——服务账号没有个人云端硬盘存储配额,上传会失败,除非使用 Google Workspace 的共享云端硬盘。

Kaggle(自动)——API 令牌设置

1. 需要一个 Kaggle 账号,并 完成手机验证(kaggle.com ▸ Settings ▸ Phone Verification)—— 否则无法开启 GPU 和 Internet。

2. 获取凭据(二选一):① kaggle.json(推荐) —— Settings ▸ API ▸ Create New Token / Create Legacy API Key,下载 kaggle.json(内含 username+key),在插件 Kaggle token 字段点 … 选这个文件;② 新式 KGAT 令牌 —— Settings ▸ API ▸ Generate new token,复制 KGAT_… 粘贴到 Kaggle token 字段,并在 Kaggle username 填你的 Kaggle 用户名。

GPU 类型的重要提醒:插件通过 Kaggle API 自动推送的 kernel 默认分到 P100(Kaggle API 无法指定 T4),而当前 PyTorch 已不支持 P100,notebook 会自动回退到 CPU(慢)。想用 GPU:打开推送出来的 kernel → 右侧 Session options ▸ Accelerator = GPU T4 x2 → Save Version;或者直接改用 Google Colab(默认就是 T4,最省心)。

百度飞桨 AI Studio / 阿里 ModelScope(仅手动)

只需一个对应平台的账号,无需任何 API 令牌。这两个平台仅支持手动模式:插件导出 input.zip + notebook;你在其网页 notebook 中把 input.zip 上传到工作目录并运行,再下载 output.zip,回插件用"Load…"按钮导入。国内可直连,适合 Colab 被墙或过慢时。

一次性环境(Setup Cloud Transfer)

自动模式(Colab/Kaggle)首次使用需点一次 Setup Cloud Transfer:用 Dragonfly 自带的 Python 建一个只含 Google Drive + Kaggle 客户端的小型虚拟环境(不含 torch,几 MB)。手动模式与 Remote Dragonfly Server 无需此步。

4. 远程 Dragonfly 服务器(第 5 个选项)

选择 Remote Dragonfly Server (via Prototype Apps) 时,插件通过 TCP 网络连接到另一台(或本机)正在运行 Prototype Labs 服务器的 Dragonfly,把 Cellpose 的训练/推理放到那台机器上执行:本机把数据(图像 + 标注)上传到服务器 → 在服务器自带的 Dragonfly Python(已内置 cellpose + torch,用它的 GPU)上训练/推理 → 把训练好的模型 / 结果 MultiROI 下载回本机并发布。无需 notebook、无需 venv、无需 pip 安装,也无需任何云账号或翻墙。 适合有一台带好显卡的机器可共享的团队。

第一步:在“服务器机器”上启动 TCP 服务器

1. 在作为服务器的那台机器上打开 Dragonfly(它需要装好 Prototype Apps / Prototype Labs)。

2. 打开 Prototype Apps ▸ Server Management...,在 TCP server 一栏点 Start。默认绑定 127.0.0.1:54321(仅本机可连)。

3. 若要从别的机器(局域网 / VPN / Tailscale)连接:在 Server Management 里把 TCP server bind(绑定地址) 下拉改成该机器的 IP —— 下拉会自动列出本机的 Tailscale 100.64.x 与局域网 IP(也可选 0.0.0.0 全部网卡)—— 再点 Start,无需设环境变量、无需重启(若服务器已在跑,先 Stop 再 Start 才会换绑定)。非本机绑定强制要求令牌:先点 Manage Access Tokens… 生成一个 token(或设 DF_SERVER_TOKEN)。别忘了在防火墙放行入站 TCP 54321。记下该机器的 IP、端口(默认 54321)与 token。

4. 多用户令牌(推荐):除了共享的 DF_SERVER_TOKEN 主令牌,还可以在 Server Management ▸ Access Tokens 里给每个人/每台终端单独生成一个 token(只存哈希,创建时显示一次),可单独启用/停用/删除。每个用户在插件的 Token 字段填自己的那个即可。

5. 确保该机器的防火墙放行了这个端口。

第二步:在本插件里连接并运行

1. 顶部 Cloud server 选择 Remote Dragonfly Server (via Prototype Apps)(模式会自动锁定为 Auto)。

2. 在 Remote server 一行填服务器的 IP 与 Port;若服务器设了令牌,则在 Token 里填相同的密钥(本机 localhost 可留空)。

3. 点 Test connection 验证连通(会显示服务器 pid / 是否在 Dragonfly 内 / 端口)。

4. 之后按左栏“训练”/右栏“推理”正常操作,点 Train on the remote server / Infer on the remote server → publish 即可;数据传输、执行、下载全部自动完成。

要点:①服务器必须是在 Dragonfly 内启动的(才有 cellpose),Test connection 会提示 in_gui;②多个用户可各用自己的 token 同时连入,但重活(训练/推理)在同一块 GPU 上仍是串行的——第 2 个人的任务会自动排队,前一个跑完再跑;③训练在服务器上跑时没有逐步进度回传,进度条会一直转,请耐心等待其完成;④点 Cancel 只能在“上传/下载”这类步骤之间生效,已在服务器上跑起来的训练要等它本轮结束;⑤令牌只是访问口令,请勿把服务器暴露到公网。

5. 安装

1. 运行安装脚本(任意 Python 3 均可):python install_cloud_cellpose_plugin.py

2. 完全重启 Dragonfly。

3. 在 Prototype Apps ▸ Cloud GPU ▸ Train/Infer Cellpose on Cloud Server... 打开插件。

6. 使用:训练(左栏)

1. 顶部选择 Cloud server 与 Mode(自动模式先按上一节配好凭据 + Setup Cloud Transfer)。

2. 选择图像 Channel + 一个已标注目标的 MultiROI(每个标签=一个目标)、方向与切片索引,点 Add pair(切片 0 = 整卷所有已标注切片)。

3. 设置基础模型、训练轮数、学习率、新模型名,点 Train on <平台>。

4. 自动模式:插件上传→打开 notebook→你 Run all→插件自动下载模型并导入本地 Models 文件夹,训练后会自动刷新并在推理栏选中该模型。

5. 手动模式:点 Export,在云端运行 notebook,下载 output.zip,再点 Load trained model 导入。

7. 使用:推理(右栏)

1. 点 Refresh,选择一个本地 模型 与一个图像 Channel。

2. 设置拼接(stitch)、flow、cellprob 阈值、直径,以及是否真三维(do_3D)。

3. 点 Infer on <平台>;自动模式会自动下载结果并把 MultiROI 发布回 Dragonfly。手动模式:运行 notebook、下载 output.zip,保持所选 Channel 不变,点 Load result & publish。

8. 常见问题排查

  • Access blocked / Error 403: access_denied(Colab 登录被拦):OAuth 同意屏幕在 Testing,且你的账号不在 Test users 里。把你的 Gmail 加为 test user,重试并在"未验证应用"页点 Advanced ▸ 继续 ▸ Allow。
  • input.zip not found under /kaggle/input:数据集没挂上。点 notebook 右上 + Add Input 添加含 input.zip 的数据集(手动模式先把 input.zip 传成 Kaggle Dataset)。Kaggle 会自动解压 zip,notebook 已兼容两种情况。
  • pip 装 cellpose 一直重试 / Temporary failure in name resolution:没开网。Settings ▸ Internet = On(需手机验证)。
  • cannot import name '_center' from numpy._core.umath:pip 改动了 host 的 numpy 造成 .py/.so 不一致。插件已改为安装 cellpose 时不动 numpy/scipy;请在干净的内核(Kaggle: Factory reset,或新开 session)里重跑。
  • CUDA error: no kernel image is available / sm_60:分到的是 P100,当前 PyTorch 不支持。改用 T4(见 Kaggle 设置)或让它回退 CPU;或用 Colab(默认 T4)。
  • Kaggle 一直 queued 不下载结果:旧版本状态读取有 bug,已修复——请重装最新插件;或手动下载 output.zip 再用 Load… 导入。
  • 训练完模型没出现在推理列表:点顶部 Refresh channels / MultiROIs / models;新版已在训练后自动刷新并选中。

9. 环境需求与注意事项

  • 云 notebook 平台(前四家)需要相应账号与网络连接;本机无需独立显卡(GPU 由云端 / 远程服务器提供)。
  • 自动模式(Colab/Kaggle)需对应凭据文件;手动模式(前四家)无需任何 API 或环境配置;Remote Dragonfly Server 只需服务器 IP/端口/令牌,不需要任何云账号或翻墙。
  • 各免费额度有上限且会变动(如 Kaggle 每周约 30 GPU 小时、Colab 单次时长限制、AI Studio 每日算力点),超大训练可能需要付费档位。凭据文件只保存在本机,并已被 .gitignore 排除。


Part II English Manual

Contents

1. Introduction

2. Cloud servers & modes

3. Accounts & credentials setup (required for Auto)

4. Remote Dragonfly Server (the 5th option)

5. Installation

6. Usage: Train (left column)

7. Usage: Infer (right column)

8. Troubleshooting

9. Requirements & notes

1. Introduction

This plugin runs both Cellpose training and inference on a cloud server's free GPU, with the same Train (left) | Infer (right) two-column layout as the built-in Cellpose plugin plus a Cloud server dropdown. Use it when the PC has no discrete GPU (or one too small for Cellpose-SAM): the PC only reads data and publishes results while the GPU-heavy compute runs in the cloud.

Why there is a browser step: free cloud notebooks have no public headless "run" API, so the GPU work always runs in a browser tab where you click Run all. The plugin generates that notebook and automates everything around it.

2. Cloud servers & modes

The top dropdown picks one of five hosts; a Mode toggle chooses Auto or Manual:

Host

Auto

Manual

Credentials

Google Colab

Yes

Yes

OAuth client_secret.json or service_account.json

Kaggle

Yes

Yes

kaggle.json file, or a KGAT_ token + username

Baidu AI Studio

-

Yes

none (Manual)

Alibaba ModelScope

-

Yes

none (Manual)

Remote Dragonfly Server (via Prototype Apps)

Yes

-

server IP + port (+ token)

  • Manual (the four notebook hosts): the plugin exports input.zip + a host-tailored notebook; you upload/run/download output.zip in the browser and load it back. No account API or setup.
  • Auto (Colab, Kaggle and Remote Dragonfly Server): the plugin uploads the data, runs the compute and downloads the result. Colab/Kaggle need a one-time account/credentials setup (next section); the Remote Dragonfly Server only needs a server IP/port/token (see its section).
  • Remote Dragonfly Server: connect to another machine (or this one) running the Prototype Labs TCP server and run training/inference on that machine's GPU. It uses the server's own Dragonfly Python (cellpose + torch already bundled), so there is no notebook, no venv, no pip install and no Google/Kaggle account or VPN. Auto-only.

3. Accounts & credentials setup (required for Auto)

Manual mode needs no credentials — skip this section. Auto mode (Colab / Kaggle) needs the accounts, API tokens and permissions set up once, as below.

Google Colab (Auto) — OAuth setup

You need a Google account. An OAuth client lets the plugin use your own Google Drive as the transfer bus.

1. Open the Google Cloud Console (console.cloud.google.com) and create/select a project.

2. Enable the Google Drive API under APIs & Services ▸ Library.

3. Configure the OAuth consent screen (newer console: Google Auth Platform ▸ Audience): User type = External; fill in the app name and support email.

4. Under Test users, add your own Gmail — this step is mandatory, otherwise sign-in is blocked with Access blocked … / Error 403: access_denied.

5. Under Credentials ▸ Create credentials ▸ OAuth client ID, choose application type Desktop app, then download client_secret.json.

6. In the plugin's Google credentials field, click … and select that client_secret.json (it should show detected: Google OAuth client).

7. The first Train/Infer on Google Colab opens a browser sign-in: sign in with your (test-user) account → Google hasn't verified this app → click Advanced ▸ Go to <app> (unsafe) → Allow Google Drive access. The token is cached, so you only do this once.

Share with (email) is unused in OAuth mode — leave it blank (it's only for the service-account path). A service_account.json also works, but is NOT recommended for a personal Gmail: service accounts have no personal Drive storage quota, so uploads fail unless you use a Google Workspace Shared Drive.

Kaggle (Auto) — API token setup

1. You need a Kaggle account with Phone Verification done (kaggle.com ▸ Settings ▸ Phone Verification) — otherwise you cannot enable GPU or Internet.

2. Get credentials (either): (1) kaggle.json (recommended) — Settings ▸ API ▸ Create New Token / Create Legacy API Key, download kaggle.json (username+key), and select it in the plugin's Kaggle token field via …; (2) the new KGAT token — Settings ▸ API ▸ Generate new token, copy the KGAT_… string into Kaggle token, and put your Kaggle username in Kaggle username.

Important GPU note: the kernel the plugin auto-pushes via the Kaggle API defaults to a P100 (the API can't request T4), and current PyTorch no longer supports the P100, so the notebook falls back to CPU (slow). For GPU speed, open the pushed kernel → Session options ▸ Accelerator = GPU T4 x2 → Save Version; or just use Google Colab (always a T4, the smoothest option).

Baidu AI Studio / Alibaba ModelScope (Manual only)

Only an account is needed — no API token. These two are Manual-only: the plugin exports input.zip + a notebook; you upload input.zip into their web notebook's working dir, run it, download output.zip, then load it back with the Load… buttons. Both are reachable in mainland China where Colab may be blocked/slow.

One-time environment (Setup Cloud Transfer)

Auto mode (Colab/Kaggle) needs a one-time Setup Cloud Transfer: it builds a tiny venv from Dragonfly's own Python holding only the Google Drive + Kaggle clients (no torch, a few MB). Manual mode and the Remote Dragonfly Server need no venv.

4. Remote Dragonfly Server (the 5th option)

With Remote Dragonfly Server (via Prototype Apps) selected, the plugin connects over TCP to another machine (or this one) that is running the Prototype Labs server inside Dragonfly, and runs Cellpose training/inference on that machine: your data (images + labels) is uploaded to the server → training/inference runs on the server's own Dragonfly Python (cellpose + torch already bundled, using its GPU) → the trained model / result MultiROI is downloaded back and published locally. No notebook, no venv, no pip install, and no cloud account or VPN. Ideal for a team with one GPU box to share.

Step 1 — start the TCP server on the SERVER machine

1. On the machine that will act as the server, open Dragonfly (it must have Prototype Apps / Prototype Labs installed).

2. Open Prototype Apps ▸ Server Management... and click Start on the TCP server row. By default it binds 127.0.0.1:54321 (this machine only).

3. To connect from another machine (LAN / VPN / Tailscale): in Server Management set the TCP server bind dropdown to that machine's IP — it auto-lists this machine's Tailscale 100.64.x and LAN IPs (or pick 0.0.0.0 for all interfaces) — then click Start. No environment variable or restart needed (if a server is already running, Stop then Start to re-bind). A non-loopback bind requires a token: click Manage Access Tokens… and Generate one first (or set DF_SERVER_TOKEN). Allow inbound TCP 54321 through the firewall. Note the IP, port (default 54321) and token.

4. Per-user tokens (recommended): besides the shared DF_SERVER_TOKEN master token, you can give each person / terminal its own token in Server Management ▸ Access Tokens (only a hash is stored; the full token is shown once), and enable / disable / delete each independently. Each user puts their own token in the plugin's Token field.

5. Make sure that port is allowed through the server machine's firewall.

Step 2 — connect and run from this plugin

1. Set Cloud server to Remote Dragonfly Server (via Prototype Apps) (the mode locks to Auto).

2. In the Remote server row, enter the server's IP and Port; if the server was started with a token, put the same secret in Token (leave blank for localhost).

3. Click Test connection to verify (it reports the server pid / whether it is inside Dragonfly / the port).

4. Then use the Train (left) / Infer (right) columns as usual and click Train on the remote server / Infer on the remote server → publish; upload, execution and download are all automatic.

Key points: (1) the server must be started inside Dragonfly (so cellpose is available) — Test connection shows in_gui; (2) multiple users can connect at once, each with their own token, but heavy jobs (train/infer) still run one at a time on the single GPU — a second user's job automatically queues; (3) while a job runs on the server there is no step-by-step progress streamed back, so the progress bar just spins until it finishes — be patient; (4) Cancel only takes effect between steps (upload/download) — a training already running on the server finishes its current run; (5) the token is just an access secret — do not expose the server to the public internet.

5. Installation

1. Run the installer (any Python 3): python install_cloud_cellpose_plugin.py

2. Fully restart Dragonfly.

3. Open it at Prototype Apps ▸ Cloud GPU ▸ Train/Infer Cellpose on Cloud Server...

6. Usage: Train (left column)

1. Pick a Cloud server and Mode at the top (for Auto, set up credentials + Setup Cloud Transfer per the section above).

2. Choose an image Channel + a MultiROI (one label per object), an orientation and slice, then Add pair (slice 0 = every labeled slice).

3. Set the base model, epochs, learning rate and model name, then Train on <host>.

4. Auto: the plugin uploads → opens the notebook → you Run all → it downloads the model and imports it into the Models folder, auto-refreshing and selecting it in the Infer column.

5. Manual: click Export, run the notebook on the host, download output.zip, then Load trained model.

7. Usage: Infer (right column)

1. Click Refresh, pick a local model and an image Channel.

2. Set the stitch, flow and cellprob thresholds, the diameter and whether to use true 3D (do_3D).

3. Click Infer on <host>; Auto downloads the result and publishes the MultiROI back into Dragonfly. Manual: run the notebook, download output.zip, keep the same Channel selected, then Load result & publish.

8. Troubleshooting

  • Access blocked / Error 403: access_denied (Colab sign-in): the OAuth consent screen is in Testing and your account isn't a Test user. Add your Gmail as a test user, retry, and click Advanced ▸ proceed ▸ Allow on the "unverified app" page.
  • input.zip not found under /kaggle/input: the dataset isn't attached. Click + Add Input (top-right) and attach the dataset with input.zip (Manual: first upload input.zip as a Kaggle Dataset). Kaggle auto-extracts zips; the notebook handles both cases.
  • pip keeps retrying / Temporary failure in name resolution: internet is off. Settings ▸ Internet = On (needs a phone-verified account).
  • cannot import name '_center' from numpy._core.umath: pip changed the host numpy, splitting its .py/.so. The plugin now installs cellpose without touching numpy/scipy; re-run in a fresh kernel (Kaggle: Factory reset, or a new session).
  • CUDA error: no kernel image is available / sm_60: you got a P100, unsupported by current PyTorch. Switch to T4 (see Kaggle setup) or let it fall back to CPU; or use Colab (T4).
  • Kaggle stuck on queued, never downloads: an old build had a status-reading bug (now fixed) — reinstall the latest plugin; or download output.zip and use Load….
  • Trained model not in the Infer list: click Refresh channels / MultiROIs / models; newer builds auto-refresh and select it after training.

9. Requirements & notes

  • The notebook hosts need the matching account and internet; no discrete GPU is required locally (the cloud host / remote server provides it).
  • Auto mode (Colab/Kaggle) needs the matching credentials file; Manual mode (the four notebook hosts) needs no API or environment setup; the Remote Dragonfly Server only needs the server IP/port/token — no cloud account or VPN.
  • Free tiers cap usage and change (e.g. Kaggle ~30 GPU-hours/week, Colab session limits, AI Studio daily points); very large jobs may need a paid tier. Credentials stay on your machine and are git-ignored.
You’ve reached the end of this manual.Explore the library →