What Is Text Generation WebUI
Text Generation WebUI is an open interface for working with language models through a browser: a chat mode, a notebook mode for free-form text, and a generation settings panel without a command line. The project supports several backends: llama.cpp for CPU and light GPU servers, and Transformers and ExLlama for full GPU inference.
The tool is handy when a model needs interactive testing, comparing answers at different temperature settings, and connecting extensions — from voice input to image generation. For a plain API without an interface, vLLM fits better, while WebUI covers the scenario of a person working with a model by hand.
Server Requirements
The minimum configuration depends on the chosen backend and model size.
| Scenario | CPU | RAM | GPU |
|---|---|---|---|
| llama.cpp, 7B model | 4 cores | 16 GB | not required |
| Transformers, 7B model | 4 cores | 16 GB | 8 GB VRAM |
| ExLlama, 13B model | 6 cores | 24 GB | 16 GB VRAM |
It is easier to estimate the exact video memory for a specific model in advance — see the article how much VRAM a model needs.
Installing on a VDS
The project ships as a plain Python repository with a dependency installation script.
sudo apt update
sudo apt install -y git python3.11-venv
git clone https://github.com/oobabooga/text-generation-webui
cd text-generation-webui
python3.11 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
Installing dependencies takes a few minutes: the script sets up PyTorch, support for the chosen backends, and a web server on Gradio.
First Launch and Loading a Model
After installation, start the server and specify a port reachable from outside so you can open the interface from your work computer.
python server.py --listen --listen-port 7860 --api
A model loads through the Model tab: GGUF files go into the models/ directory, after which the model shows up in the dropdown list. The --api flag also opens an OpenAI-compatible endpoint, so WebUI can serve as a backend for third-party applications, not just a browser chat.
Useful Extensions and Modes
WebUI covers more than the chat mode — extensions handle most working scenarios:
- Notebook mode — free-form text with no assistant or user roles, handy for continuation and prompt experiments.
- The superbooga extension adds simple document search right inside the dialog.
- Character cards set the model's personality and reply style through a separate JSON profile.
- The built-in API mode lets you connect WebUI to third-party scripts the same way as OpenAI API — though for production load a separate inference server like llama.cpp directly is more reliable.
What to Do About Common Problems
Most failures are fixed by checking dependency versions or launch parameters:
- The interface does not open from outside — check the
--listenflag and the firewall rule for port 7860. - The model fails to load with a memory error — lower the context length in the model settings or pick a lighter quantization.
- The ExLlama backend cannot find the CUDA library — reinstall the PyTorch version listed in the GPU requirements file.
- Answers cut off mid-sentence — increase the max_new_tokens parameter on the generation tab.
Once configured, WebUI turns into a convenient workbench: you can quickly switch models, compare quantization levels, and share access with colleagues on the same VDS.