Managed Local AI on your own GPU

Your own private AI — like ChatGPT, but it runs on your server and your data never leaves the house. We build it on Ollama (the open-source app that runs AI models on your own machine), keep it updated, and show you real test results before you pay for performance.

  • Private AI built on Ollama, the open-source standard
  • Open-source hosting managed with CyberPanel
  • Domains, support, and clear ways to get started

Runs on Ollama with Open WebUI, a private ChatGPT-style chat for your team. Outside AI services are used only if you ask for them.

Managed Local AI

Your own private AI running on your GPU, with Open WebUI — a ChatGPT-style chat window for your team.

  • Ollama standard
  • Open WebUI default
  • No third-party AI API required by default
  • Managed setup and updates
Explore Local AI →

AI Apps

Private document search, team chat, workflow automation, and creative image tools — all open-source and managed by us.

  • AnythingLLM or LibreChat
  • Flowise and n8n options
  • ComfyUI for creative GPU workflows
  • vLLM only as optional advanced layer
Compare AI Apps →

GPU Infrastructure

The right GPU server for running AI, generating images, and private automation — sized to what you actually need.

  • Dedicated setup
  • Storage and backup planning
  • Monitoring and maintenance
  • Benchmark before performance promises
Plan GPU Stack →

Open Source Hosting

Managed open-source hosting with CyberPanel, domains, SSL, DNS, and human support.

  • CyberPanel control panel
  • WordPress, Nextcloud, Matomo and more
  • No cPanel or CentOS claims
  • Built for long-term maintenance
View Hosting Options →

Domains

Domain registration, renewal, transfer guidance, and DNS support for open-source projects and teams.

  • Popular TLDs with USD pricing
  • Renewal notes shown clearly
  • Transfers reviewed by registry rules
  • DNS basics included
Check Domains →

Your private AI, set up and kept running

We install it, secure it, keep it updated, and explain how it works. By default your AI runs entirely on your own server, with no outside AI service needed.

Default UI

Open WebUI

The chat window your team uses in the browser — like ChatGPT, but private to your server.

Knowledge

AnythingLLM

Lets your AI answer from your own documents, when that fits what you need.

Advanced

vLLM optional

An advanced speed option we only add after testing it on your model and GPU.

Connect your own tools to your private AI

Many teams don't need a new chatbot first. They need their existing software to talk to a private AI, with access kept under control. We test that connection before anyone relies on it.

API bridge

OpenAI-compatible local endpoint, tested before launch

Your private AI can speak the same "language" as the OpenAI API, so your existing scripts and apps can point to it instead — once we've checked it's ready.

Open WebUI roles and SSO scope

Who can see and do what

For teams, we set up user roles (RBAC), single sign-on (SSO/OIDC), API keys, and per-model permissions before anyone goes live.

Serving trial

Extra speed only after fit checks

There are advanced ways to make models smaller and faster on the GPU (AWQ, GPTQ, GGUF, FP8). We treat them as an optional test track, never an upfront promise.

Our promise: we test the connection on your own server and show you the results before you rely on it in production.

Private AI agents, tested safely first

An AI "agent" can take actions for you, not just chat. That is powerful, but it adds risk. Our Business Secure plan tests each agent on a fixed set of real tasks first, instead of promising a hands-off robot assistant.

Reasoning fit

gpt-oss 20B as a benchmark candidate

OpenAI positions gpt-oss models for reasoning and tool-use workflows, while Ollama lists the 20B path around 14 GB with a 128K context window (how much text it can read at once). On this server class, that remains a measured Business Secure trial.

OpenAI gpt-oss-20b model card →Ollama gpt-oss →
Guardrail fit

gpt-oss-safeguard 20B policy benchmark

OpenAI describes gpt-oss-safeguard as open safety reasoning models for custom safety policies. Treat the 20B model as a measured policy-classification benchmark for local agent workflows, not an automatic moderation guarantee.

OpenAI safeguard announcement →Hugging Face safeguard-20b →
Tool boundary

Tools are treated as privileged access

Open WebUI documents that tool access can execute arbitrary Python code. We scope tool permissions, API keys, and provider credentials before any agent workflow touches customer data or business systems.

Open WebUI permissions →
Proof

Fixed-task audit before rollout

The benchmark uses a fixed task set, allowed tools, blocked tools, sample data, logging expectations, and rollback notes. Production use waits for GPU/Ollama health and a repeatable target-workflow smoke test.

Our promise: we prove an agent works safely on a fixed set of real tasks, on your own server, before we ever call it production-ready.

First we check it works, then we pick your model

Local AI changes fast. We turn that uncertainty into a short, practical set of checks before you commit to a bigger rollout.

Step 1

Runtime check

We verify GPU drivers, CUDA visibility, Ollama or container service health, storage, backups, and secure access before model work starts.

Step 2

20 GB model shortlist

For RTX 4000 Ada class systems, standard fit checks start with quantized 8B/14B models such as Qwen3; larger 30B/32B paths remain experimental and we document the latency, context, and concurrency tradeoffs.

Step 3

Benchmark report

Performance claims are made after testing the target model, quantization, context length, and user workflow. If vLLM is useful, it stays an optional advanced layer.

Which AI models fit a 20 GB card right now

Checked against current sources and the live runtime on 2026-08-02. Start standard fit checks with Qwen3 8B for latency/concurrency or quantized Qwen3 14B for richer private assistant and RAG work; 30B/32B variants remain experimental on 20 GB. These are candidates, not guaranteed speeds.

Documents

Qwen3-VL + Docling baseline

Qwen3-VL 8B is a current OCR and document-structure candidate; Docling/OCR gives a deterministic parsing baseline for PDFs, tables, reading order, and field-level checks.

Qwen3-VL-8B model card →
Visual RAG

Qwen3-VL Embedding + Reranker trial

Qwen3-VL-Embedding-8B and Qwen3-VL-Reranker-8B are current multimodal retrieval candidates for text, images, screenshots, videos, and mixed documents. On 20 GB systems, scope them as a measured visual-RAG benchmark with corpus size, vector storage, latency, and fallback limits.

Qwen3-VL-Embedding-8B card →Qwen3-VL-Reranker-8B card →
Code

Qwen3-Coder 30B benchmark

Ollama lists a 19 GB model package with a large context window; that nearly fills 20 GB before runtime and KV-cache overhead. Treat it as benchmark-only with strict context and concurrency limits, with an 8B/14B fallback.

Ollama Qwen3-Coder listing →
Assistant

Gemma 4 E4B/26B trial

Gemma 4 E4B is the low-memory assistant and multimodal candidate. Gemma 4 26B and 31B are benchmark-only on 20 GB because Ollama lists 18 GB and 20 GB footprints; concurrency, context, and MTP latency gains stay gated by local smoke tests.

Google Gemma 4 announcement →Gemma 4 MTP drafters →
Reasoning

gpt-oss 20B private reasoning benchmark

OpenAI positions gpt-oss-20b for local and specialized reasoning use, and Ollama lists the 20B path around 14 GB with 128K context. Treat it as a strong Business Secure benchmark candidate with strict context, tool-use, safety, and latency checks before rollout.

OpenAI gpt-oss-20b card →OpenAI model card PDF →

Our promise: we make no speed or performance claim until the drivers, Ollama, and your chosen model all pass tests on your real server.

Four ready-made private-AI packages

Four practical packages for a 20 GB GPU. Each one starts with a real test on your server (a "smoke test") before we make any performance claim.

Team knowledge

Private RAG with sources

An AI that answers using your own documents and shows its sources — this is called RAG. Good for internal files, support archives, policies, and project knowledge, where every answer should point back to where it came from.

  • Open WebUI or lightweight team UI
  • Qwen3 or Gemma 4 chat candidate
  • EmbeddingGemma or Qwen embedding trial

Smoke test: ingest a representative corpus, ask fixed benchmark questions, require cited answers, and record VRAM, latency, and miss behavior.

Start €431 RAG quality audit →
Documents

Local PDF and image extraction

For invoices, forms, screenshots, and operational documents that need local extraction support without sending files to a cloud AI API by default.

  • Qwen3-VL/Qwen2.5-VL benchmark
  • Docling/OCR fallback for hard scans
  • Field-level error notes

Smoke test: run real sample pages against expected fields, measure false positives, unsupported layouts, throughput, and VRAM peak.

€431 one-time pilot scope: up to 10 representative files or images (maximum 50 pages total) for one defined extraction workflow. Extra pages, fields, unsupported layouts, or rework are quoted separately; there is no page-volume SLA.

Read document intake benchmark →Request document workflow trial →
Visual search

Visual RAG and evidence search

For teams that need to retrieve answers from screenshots, diagrams, scanned pages, product images, or short video captures with evidence links instead of plain text-only RAG.

  • Qwen3-VL-Embedding 8B recall trial
  • Qwen3-VL-Reranker 8B ranking trial
  • Source thumbnails and miss analysis

Smoke test: index a small mixed-media corpus, ask fixed visual-search questions, record top-k misses, reranker gains, storage size, VRAM, and latency before rollout.

Start €431 RAG quality audit →
Audio

Local transcription and meeting notes

For interviews, internal meetings, and support recordings where private audio handling and predictable operations matter more than a generic SaaS workflow.

  • Whisper or faster-whisper benchmark
  • English and German sample set
  • Optional local summary pass

Smoke test: transcribe representative 5 to 30 minute files, record runtime factor, language quality, segmentation limits, and GPU use.

Scope transcription benchmark →

Our promise: these packages are sold as setup and ongoing management. Any live performance claim waits until it is tested on your actual server.

Current visual/document evidence (2026-08-02): qwen3-vl:8b passed a fixed public synthetic EZOS visual/document benchmark: invoice OCR, Team RAG screenshot interpretation, and local runtime-gate panel reading, 3/3 cases. This proves the benchmark path on synthetic public sample images, not customer PDF, screenshot, or production corpus performance.

RAG Retrieval Quality Audit

A benchmark-first entry into Team RAG. For teams that want an AI to search their own documents but aren't yet sure their files, their questions, and their access rules are ready for a full monthly rollout.

€431 one-time. A bounded audit for up to 25 representative files (maximum 200 pages total) and up to 15 fixed questions, with citation checks and a go/no-go report before recurring Team RAG.

Upgrade path: when the recurring Team RAG scope is approved within 30 days, the €431 audit fee is credited against its setup fee and confirmed before provisioning. Extra files, pages, questions, or rework are quoted separately.

Scope

Representative corpus first

We start with a limited set of real documents, screenshots, policies, tickets, or manuals and turn them into a fixed retrieval benchmark instead of ingesting everything blindly.

Models

Qwen3 small-to-medium retrieval path

Qwen3-Embedding 0.6B/4B and Qwen3-Reranker 0.6B/4B are the default audit candidates. The 8B path stays optional when corpus size, latency, and the 20 GB VRAM budget justify it.

Proof

Answer quality before rollout

The result is a short report with top-k misses, citation quality, reranker gains, storage notes, privacy boundaries, and a clear decision: fix sources, run a pilot, or move to Team RAG.

Current local evidence (2026-08-02): qwen3-embedding:0.6b + gpt-oss:20b passed a fixed public EZOS Team RAG/agent benchmark: 5 retrieval cases, 0.8 top-1 retrieval accuracy, and 3/3 grounded answer plus agent-boundary cases. This proves the benchmark path on public sample data, not customer-corpus production performance.

Our promise: this audit gives you real evidence of search quality and a clear next step. It makes no live-performance claim until the GPU, Ollama, and your chosen models pass tests on your own server.

Plans built around real work

Use this page to pick the right starting point, then confirm the exact setup, renewal terms, and any test notes at checkout.

BYO server

BYO Server Management

from $295.10/mo

Already have your own GPU server or rented one? We manage Ollama, Open WebUI, updates, and monitoring on it. (BYO means bring your own server.)

  • Driver and Ollama readiness check
  • Open WebUI installation
  • Hosting costs paid directly by you
Team knowledge

Team RAG

from $999.60/mo

For teams that need document-assisted local AI, curated embeddings, user roles, and clearer knowledge workflows.

  • Knowledge/RAG setup
  • Embedding and model guidance
  • Prioritized support
Controlled rollout

Business Secure

from $1,499.90/mo

For production business use where security, audit preparation, update windows, and support scope matter.

  • Security hardening
  • OIDC/SSO preparation
  • Monthly review and change windows

Before we make any live-AI claim, we check the GPU drivers and confirm Ollama is healthy. A 20 GB RTX 4000 Ada card is great for small-to-medium models, not the very largest frontier models.

Public starting prices are shown in USD and were checked against the secure cart on 2026-08-03. Checkout defaults to EUR, can switch currency, and remains authoritative for setup fees, term, renewal, taxes, and final total.

Not ready for checkout? Ask first.

Goes to sales@ezoshosting.com. Reply by email.