How to Run Ollama on a VPS: Self-Hosted LLM Inference
Run Ollama on a VPS for self-hosted LLM inference: CPU vs GPU sizing, systemd tuning, nginx auth, the OpenAI-compatible API and when to move to vLLM.
Read MoreGPU servers, LLM inference and machine learning workloads on dedicated infrastructure. 8 articles from the MassiveGRID engineering teams.
Run Ollama on a VPS for self-hosted LLM inference: CPU vs GPU sizing, systemd tuning, nginx auth, the OpenAI-compatible API and when to move to vLLM.
Read MoreHow much VRAM a local LLM really needs: weight sizes by quantization, the KV cache formula, concurrency maths and which GPU fits which model.
Read MoreServe an OpenAI-compatible LLM API with vLLM: continuous batching, PagedAttention, the flags that matter, systemd setup, nginx and throughput benchmarking.
Read MoreA100, H100 and RTX 6000 Ada compared for AI: memory bandwidth vs compute, FP8, NVLink, hourly vs monthly cost, and which card suits each workload.
Read MoreEuropean GPU hosting compared: GDPR transfers and the CLOUD Act, where transatlantic latency actually hurts, egress costs, and choosing Frankfurt or London.
Read MoreRunning AI and ML Workloads on a VPS: Training vs Inference: Two Very Different Workloads and AI/ML Workloads That Run Well on a VPS.
Read MoreBest VPS for n8n AI Agents: n8n as an AI Agent Platform, VPS Requirements for AI Workloads and Sizing Tiers: API-Only, Hybrid, and Local LLM.
Read MoreNVIDIA H100 vs H200: The NVIDIA H100 and H200 are both high-performance GPUs based on NVIDIA’s Hopper architecture, designed for AI, high-performance.
Read More