Skip to main content

Text Generation WebUI: A Web Interface for LLMs

AI Agents on VPS · 29.09.2026

What Is Text Generation WebUI

Text Generation WebUI is an open interface for working with language models through a browser: a chat mode, a notebook mode for free-form text, and a generation settings panel without a command line. The project supports several backends: llama.cpp for CPU and light GPU servers, and Transformers and ExLlama for full GPU inference.

The tool is handy when a model needs interactive testing, comparing answers at different temperature settings, and connecting extensions — from voice input to image generation. For a plain API without an interface, vLLM fits better, while WebUI covers the scenario of a person working with a model by hand.

Server Requirements

The minimum configuration depends on the chosen backend and model size.

ScenarioCPURAMGPU
llama.cpp, 7B model4 cores16 GBnot required
Transformers, 7B model4 cores16 GB8 GB VRAM
ExLlama, 13B model6 cores24 GB16 GB VRAM

It is easier to estimate the exact video memory for a specific model in advance — see the article how much VRAM a model needs.

Installing on a VDS

The project ships as a plain Python repository with a dependency installation script.

sudo apt update
sudo apt install -y git python3.11-venv
git clone https://github.com/oobabooga/text-generation-webui
cd text-generation-webui
python3.11 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Installing dependencies takes a few minutes: the script sets up PyTorch, support for the chosen backends, and a web server on Gradio.

First Launch and Loading a Model

After installation, start the server and specify a port reachable from outside so you can open the interface from your work computer.

python server.py --listen --listen-port 7860 --api

A model loads through the Model tab: GGUF files go into the models/ directory, after which the model shows up in the dropdown list. The --api flag also opens an OpenAI-compatible endpoint, so WebUI can serve as a backend for third-party applications, not just a browser chat.

Useful Extensions and Modes

WebUI covers more than the chat mode — extensions handle most working scenarios:

  • Notebook mode — free-form text with no assistant or user roles, handy for continuation and prompt experiments.
  • The superbooga extension adds simple document search right inside the dialog.
  • Character cards set the model's personality and reply style through a separate JSON profile.
  • The built-in API mode lets you connect WebUI to third-party scripts the same way as OpenAI API — though for production load a separate inference server like llama.cpp directly is more reliable.

What to Do About Common Problems

Most failures are fixed by checking dependency versions or launch parameters:

  • The interface does not open from outside — check the --listen flag and the firewall rule for port 7860.
  • The model fails to load with a memory error — lower the context length in the model settings or pick a lighter quantization.
  • The ExLlama backend cannot find the CUDA library — reinstall the PyTorch version listed in the GPU requirements file.
  • Answers cut off mid-sentence — increase the max_new_tokens parameter on the generation tab.

Once configured, WebUI turns into a convenient workbench: you can quickly switch models, compare quantization levels, and share access with colleagues on the same VDS.

← Back to Knowledge Base Ask Support