Skip to main content

ComfyUI on a Server: Building Image Generation Pipelines

AI Agents on VPS · 29.09.2026

A plain Stable Diffusion web interface gives you one form: prompt, size, number of steps. ComfyUI replaces the form with a node graph — you can assemble a pipeline from checkpoint loading, upscaling, inpainting, and ControlNet in any order and reuse it as a template. Let's see how to set up ComfyUI on a GPU VDS and run pipelines headlessly, through the API.

How ComfyUI differs from a regular interface

The basic Stable Diffusion setup, described in the article on image generation on a server, is fine for simple text-to-image tasks. ComfyUI is built as a visual graph editor: every step — loading a checkpoint, encoding the prompt, sampling, VAE decoding — is a separate node with its own inputs and outputs. By connecting nodes, you can assemble a complex pipeline: for example, generation, two-pass upscaling, and a ControlNet overlay in a single run.

The main advantage for a server is that a finished graph is saved as JSON and runs programmatically through the API, with no browser open and no manual clicking through buttons.

Installing on a GPU VDS

ComfyUI needs Python 3.10+, a CUDA-compatible GPU driver, and about 10 GB of disk space for the program itself, not counting model weights.

git clone https://github.com/comfyanonymous/ComfyUI.git
cd ComfyUI
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
python3 main.py --listen 0.0.0.0 --port 8188

Model checkpoints go into models/checkpoints/, LoRA adapters into models/loras/, and ControlNet models into models/controlnet/. After the first launch, the interface is available at the server's address on port 8188; for a production environment, close that port with a firewall and allow access only through a reverse proxy with basic authentication.

Basics: nodes, the graph, and checkpoints

A minimal graph for generating an image consists of five nodes: CheckpointLoader loads the model weights, CLIPTextEncode encodes the positive and negative prompt, KSampler runs diffusion, VAEDecode turns the latent into pixels, and SaveImage saves the result. Each node is configured separately: KSampler has steps, cfg, and sampler_name parameters that a regular web interface hides behind a single form.

Finished graphs can be saved and reused: a workflow file in JSON format includes every node, its parameters, and the connections between them.

Running a pipeline through the API

ComfyUI accepts a finished graph through a REST API — a key feature for automation without an open interface. A generation request is sent with POST, with the body as a JSON graph:

curl -X POST http://127.0.0.1:8188/prompt \
  -H "Content-Type: application/json" \
  -d @workflow.json

The response contains a prompt_id you can use to poll the queue via GET /history/{prompt_id} and get the path to the finished file once rendering completes. This approach fits well with a bot or website integration: the user sends a prompt, the backend inserts it into a graph template, and calls the ComfyUI API.

Extensions: ControlNet, LoRA, and custom nodes

The ComfyUI community releases hundreds of custom nodes — they are installed through the ComfyUI-Manager extension or manually into the custom_nodes/ folder. The most common additions are ControlNet for controlling composition from a sketch or a depth map, IPAdapter for transferring style from a reference image, and upscaling nodes built on Real-ESRGAN.

After installing a new node, ComfyUI needs a restart — nodes register at server startup and are not loaded on the fly.

How many resources the server needs

ModelResolutionVRAMTime per frame (RTX class)
SD 1.5512x5124-6 GB2-4 sec
SDXL1024x10248-12 GB6-10 sec
SDXL + ControlNet1024x102412-16 GB10-15 sec
SDXL + 2x upscale2048x204816-20 GB20-30 sec

For the formula and tables to size VRAM precisely for a specific model, see the article how much VRAM a model needs — the principle overlaps between language models and diffusion ones: the higher the resolution and the longer the pipeline, the bigger the memory margin needed for intermediate tensors.

Checklist before going to production

Before opening ComfyUI to outside users, close port 8188 to direct access, limit the generation queue per user, and set up monitoring of the queue and VRAM so you notice load growth in time. It helps to connect general AI service monitoring on a VDS and track queue length as a separate metric.

  • Hide the ComfyUI port behind a reverse proxy with authentication, never expose it directly to the internet.
  • Save working graphs as JSON templates and substitute parameters programmatically.
  • Limit the generation queue per user so one client cannot occupy the whole GPU.
  • Budget VRAM with a margin for ControlNet and upscaling, not just for basic generation.
← Back to Knowledge Base Ask Support