Skip to main content
🤖

AI Agents on VPS

34 articles in this section

This section covers running AI tools on your own server — without relying on someone else's API and with full control over your data. It starts with what AI agents are and why you'd run them on a VDS, installing Ollama to run local LLMs, the Open WebUI interface on top of it, and how engines like vLLM, llama.cpp, and Text Generation WebUI compare on different hardware.

From there it moves to building your own systems: agents built with LangChain, multi-agent setups through CrewAI and AutoGen, no-code automation in n8n, the visual builder Flowise, and RAG for answering questions over your own documents using vector databases like Chroma, Qdrant, or pgvector inside PostgreSQL. It also covers connecting external APIs — OpenAI, Claude, and Gemini — along with their cost and how they compare.

There are practical scenarios too: a Telegram bot built on LangChain and Ollama, a PHP-and-JavaScript chat widget for a website, speech transcription with Whisper, image generation with Stable Diffusion and ComfyUI, sizing VRAM for a given model, and protecting agents from prompt injection.

  • Local LLMs: Ollama, vLLM, llama.cpp
  • Agents and RAG: LangChain, CrewAI, vector databases
  • No-code automation: n8n and Flowise
  • Monitoring and securing AI agents
AI Agents: What They Are and Why You Need Them on a VDS AI agents are autonomous programs built on LLMs. Learn how to deploy an AI agent on a VDS server and which frameworks to choose. Installing Ollama on a VDS: Running Local LLMs Ollama lets you run Llama 3, Mistral, and Phi-3 on a VDS. A step-by-step guide to installation, configuration, and running your first model. Open WebUI: A Web Interface for Ollama on a VDS Open WebUI is a ChatGPT-like interface for local Ollama models. Install via Docker, configure Nginx, and enable multi-user mode. LangChain on a VDS: Building AI Agents in Python LangChain is the leading framework for AI agents in Python. Install it on a VDS, connect to Ollama, build chains, and create your first agent with tools. n8n on a VDS: No-Code Automation with AI Agents n8n is a powerful automation tool with AI agent support. Installing it on a VDS, connecting to Ollama, and building your first AI workflow. Flowise on a VDS: a visual builder for AI agents Flowise is a drag-and-drop platform for building AI agents and LLM pipelines. Installing on a VDS, connecting Ollama, and ready-made chatbot templates. RAG on a VDS: AI Answers Based on Your Documents RAG (Retrieval-Augmented Generation) lets an LLM answer questions about your documents. A complete tutorial using LangChain, Chroma, and Ollama on a VDS. AI Telegram Bot on a VDS with LangChain and Ollama Build an AI Telegram bot with dialogue memory on a VDS using Python, LangChain, and a local Ollama model. Step-by-step guide with full code. Chroma DB on a VDS: Vector Database for RAG Chroma is a popular vector database for RAG systems. Installation on a VDS, embedded and server modes, integration with LangChain. Qdrant on a VDS: A High-Performance Vector Database Qdrant is a production-ready vector database written in Rust for billions of vectors. Installation on a VDS, Python client, and LangChain integration. OpenAI API on a VDS: Connection and Cost Optimization Connecting the OpenAI API to a VDS, a mixed Ollama+OpenAI strategy, and a LiteLLM proxy for your team. Save up to 90% on AI requests. CrewAI on a VDS: a Team of AI Agents for Complex Tasks CrewAI lets you build teams of specialized AI agents. Installation on a VDS with Ollama, creating agents with roles and tasks. LocalAI on a VDS: an OpenAI-compatible server LocalAI is a full replacement for the OpenAI API with local models. Supports text, images, and speech. Installation on a VDS via Docker. Monitoring AI Agents on a VPS: Langfuse and Metrics Monitoring AI agents in production: Langfuse for LLM tracing, token and latency metrics, LangChain integration on a VPS. AutoGen on a VDS: Microsoft's Multi-Agent Framework AutoGen is Microsoft's framework for multi-agent systems. Agents write and run code and talk to each other. Installing it on a VDS with Ollama. ChatGPT API: Connecting and Using the OpenAI API Connecting OpenAI ChatGPT API: API key setup, GPT-4o requests, streaming, PHP and Python examples, 2025 model pricing. Claude API (Anthropic): Integrating with a VPS Claude API by Anthropic: setup, PHP and Python examples, vision, comparison with ChatGPT GPT-4o, pricing for Claude 3.5 Sonnet and Haiku. Stable Diffusion on VPS: Image Generation on Your Server Deploying Stable Diffusion WebUI on GPU VPS. Installation, API, Python integration, ComfyUI, GPU server selection. OpenAI Whisper: Audio and Video Transcription on a VPS Local audio transcription with OpenAI Whisper on a VPS: installation, Python API, FastAPI server, Faster-Whisper 4x speedup, Russian and Ukrainian support. Google Gemini API: Setup and Usage on a VPS Google Gemini API: connecting, Python and PHP examples, 1M-token context, multimodal image analysis, and Gemini 1.5 Flash and Pro 2025 pricing. MCP Server: Connecting Tools to Claude AI Model Context Protocol (MCP): Claude Desktop setup, filesystem and DB connection, building custom Python MCP server for hosting automation. AI Chatbot for a Website with ChatGPT: PHP Widget Build an AI support chatbot for your website with the ChatGPT API: PHP backend, JavaScript widget, rate limiting, prompt system, VPS deployment. Ollama: Local LLMs on a VPS Without API Keys Run LLaMA 3, Mistral, and Gemma on a VPS with Ollama: installation, REST API, Python integration, model comparison, and data privacy without external services. AI API Comparison 2025: OpenAI, Claude, Gemini Detailed LLM API comparison 2025: GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, Groq, Ollama. Pricing, speed, context windows, task-specific selection. Setting Up vLLM on a Server for Fast LLM Inference Setting up vLLM on a GPU server: hardware requirements, installation steps, launching an OpenAI-compatible API, and fixing out-of-memory errors. llama.cpp and GGUF: Running Language Models on CPU How to run a language model on CPU with llama.cpp: the GGUF format, quantization levels, building the binary, and the first launch without a GPU. Text Generation WebUI: A Web Interface for LLMs Installing Text Generation WebUI on a VDS: connecting the llama.cpp and Transformers backends, loading models, and chatting through a browser. Semantic Search for a Website Using Embeddings How to build semantic search for a website using embeddings: picking a model, generating vectors, storing them, and returning relevant articles. pgvector: Vector Search Directly in PostgreSQL Installing pgvector for PostgreSQL: creating a vector column, the HNSW index, nearest-neighbor search, and common setup mistakes on a VDS. Vector Databases Compared: Qdrant, Weaviate and Milvus Vector database for RAG and search: how Qdrant, Weaviate and Milvus differ and which to run on a VDS. How Much VRAM a Language Model Needs: Sizing Guide How to calculate VRAM for a specific LLM: the formula, quantization tables, and GPU picks for a VDS. Real-Time Streaming Transcription with Whisper on a Server Streaming transcription with Whisper: audio chunking, VAD, faster-whisper, and a WebSocket server on a VDS. ComfyUI on a Server: Building Image Generation Pipelines Installing ComfyUI on a GPU VDS: nodes, ready-made workflows, running via API, and sizing VRAM. AI Agent Security: Defending Against Prompt Injection How to protect an AI agent from prompt injection: tool isolation, an action allowlist, and monitoring on a VDS.