Managed Local AI
Your own private AI running on your GPU, with Open WebUI — a ChatGPT-style chat window for your team.
- Ollama standard
- Open WebUI default
- No third-party AI API required by default
- Managed setup and updates
Your own private AI — like ChatGPT, but it runs on your server and your data never leaves the house. We build it on Ollama (the open-source app that runs AI models on your own machine), keep it updated, and show you real test results before you pay for performance.
Runs on Ollama with Open WebUI, a private ChatGPT-style chat for your team. Outside AI services are used only if you ask for them.
Your own private AI running on your GPU, with Open WebUI — a ChatGPT-style chat window for your team.
Private document search, team chat, workflow automation, and creative image tools — all open-source and managed by us.
The right GPU server for running AI, generating images, and private automation — sized to what you actually need.
Managed open-source hosting with CyberPanel, domains, SSL, DNS, and human support.
Domain registration, renewal, transfer guidance, and DNS support for open-source projects and teams.
We install it, secure it, keep it updated, and explain how it works. By default your AI runs entirely on your own server, with no outside AI service needed.
The chat window your team uses in the browser — like ChatGPT, but private to your server.
Lets your AI answer from your own documents, when that fits what you need.
An advanced speed option we only add after testing it on your model and GPU.
Many teams don't need a new chatbot first. They need their existing software to talk to a private AI, with access kept under control. We test that connection before anyone relies on it.
Your private AI can speak the same "language" as the OpenAI API, so your existing scripts and apps can point to it instead — once we've checked it's ready.
For teams, we set up user roles (RBAC), single sign-on (SSO/OIDC), API keys, and per-model permissions before anyone goes live.
There are advanced ways to make models smaller and faster on the GPU (AWQ, GPTQ, GGUF, FP8). We treat them as an optional test track, never an upfront promise.
Our promise: we test the connection on your own server and show you the results before you rely on it in production.
An AI "agent" can take actions for you, not just chat. That is powerful, but it adds risk. Our Business Secure plan tests each agent on a fixed set of real tasks first, instead of promising a hands-off robot assistant.
OpenAI positions gpt-oss models for reasoning and tool-use workflows, while Ollama lists the 20B path around 14 GB with a 128K context window (how much text it can read at once). On this server class, that remains a measured Business Secure trial.
OpenAI gpt-oss-20b model card →Ollama gpt-oss →OpenAI describes gpt-oss-safeguard as open safety reasoning models for custom safety policies. Treat the 20B model as a measured policy-classification benchmark for local agent workflows, not an automatic moderation guarantee.
OpenAI safeguard announcement →Hugging Face safeguard-20b →Open WebUI documents that tool access can execute arbitrary Python code. We scope tool permissions, API keys, and provider credentials before any agent workflow touches customer data or business systems.
Open WebUI permissions →The benchmark uses a fixed task set, allowed tools, blocked tools, sample data, logging expectations, and rollback notes. Production use waits for GPU/Ollama health and a repeatable target-workflow smoke test.
Our promise: we prove an agent works safely on a fixed set of real tasks, on your own server, before we ever call it production-ready.
Local AI changes fast. We turn that uncertainty into a short, practical set of checks before you commit to a bigger rollout.
We verify GPU drivers, CUDA visibility, Ollama or container service health, storage, backups, and secure access before model work starts.
For RTX 4000 Ada class systems, standard fit checks start with quantized 8B/14B models such as Qwen3; larger 30B/32B paths remain experimental and we document the latency, context, and concurrency tradeoffs.
Performance claims are made after testing the target model, quantization, context length, and user workflow. If vLLM is useful, it stays an optional advanced layer.
Checked against current sources and the live runtime on 2026-08-02. Start standard fit checks with Qwen3 8B for latency/concurrency or quantized Qwen3 14B for richer private assistant and RAG work; 30B/32B variants remain experimental on 20 GB. These are candidates, not guaranteed speeds.
Qwen3-Embedding and Qwen3-Reranker 0.6B/4B/8B cover multilingual retrieval, code search, and source ranking. Test corpus quality, storage, latency, and answer citations before rollout.
Qwen3 Embedding announcement →Qwen3-Embedding-4B card →Qwen3-VL 8B is a current OCR and document-structure candidate; Docling/OCR gives a deterministic parsing baseline for PDFs, tables, reading order, and field-level checks.
Qwen3-VL-8B model card →Qwen3-VL-Embedding-8B and Qwen3-VL-Reranker-8B are current multimodal retrieval candidates for text, images, screenshots, videos, and mixed documents. On 20 GB systems, scope them as a measured visual-RAG benchmark with corpus size, vector storage, latency, and fallback limits.
Qwen3-VL-Embedding-8B card →Qwen3-VL-Reranker-8B card →Ollama lists a 19 GB model package with a large context window; that nearly fills 20 GB before runtime and KV-cache overhead. Treat it as benchmark-only with strict context and concurrency limits, with an 8B/14B fallback.
Ollama Qwen3-Coder listing →Gemma 4 E4B is the low-memory assistant and multimodal candidate. Gemma 4 26B and 31B are benchmark-only on 20 GB because Ollama lists 18 GB and 20 GB footprints; concurrency, context, and MTP latency gains stay gated by local smoke tests.
Google Gemma 4 announcement →Gemma 4 MTP drafters →OpenAI positions gpt-oss-20b for local and specialized reasoning use, and Ollama lists the 20B path around 14 GB with 128K context. Treat it as a strong Business Secure benchmark candidate with strict context, tool-use, safety, and latency checks before rollout.
OpenAI gpt-oss-20b card →OpenAI model card PDF →Our promise: we make no speed or performance claim until the drivers, Ollama, and your chosen model all pass tests on your real server.
Four practical packages for a 20 GB GPU. Each one starts with a real test on your server (a "smoke test") before we make any performance claim.
An AI that answers using your own documents and shows its sources — this is called RAG. Good for internal files, support archives, policies, and project knowledge, where every answer should point back to where it came from.
Smoke test: ingest a representative corpus, ask fixed benchmark questions, require cited answers, and record VRAM, latency, and miss behavior.
Start €431 RAG quality audit →For invoices, forms, screenshots, and operational documents that need local extraction support without sending files to a cloud AI API by default.
Smoke test: run real sample pages against expected fields, measure false positives, unsupported layouts, throughput, and VRAM peak.
€431 one-time pilot scope: up to 10 representative files or images (maximum 50 pages total) for one defined extraction workflow. Extra pages, fields, unsupported layouts, or rework are quoted separately; there is no page-volume SLA.
Read document intake benchmark →Request document workflow trial →For teams that need to retrieve answers from screenshots, diagrams, scanned pages, product images, or short video captures with evidence links instead of plain text-only RAG.
Smoke test: index a small mixed-media corpus, ask fixed visual-search questions, record top-k misses, reranker gains, storage size, VRAM, and latency before rollout.
Start €431 RAG quality audit →For interviews, internal meetings, and support recordings where private audio handling and predictable operations matter more than a generic SaaS workflow.
Smoke test: transcribe representative 5 to 30 minute files, record runtime factor, language quality, segmentation limits, and GPU use.
Scope transcription benchmark →Our promise: these packages are sold as setup and ongoing management. Any live performance claim waits until it is tested on your actual server.
Current visual/document evidence (2026-08-02): qwen3-vl:8b passed a fixed public synthetic EZOS visual/document benchmark: invoice OCR, Team RAG screenshot interpretation, and local runtime-gate panel reading, 3/3 cases. This proves the benchmark path on synthetic public sample images, not customer PDF, screenshot, or production corpus performance.
A benchmark-first entry into Team RAG. For teams that want an AI to search their own documents but aren't yet sure their files, their questions, and their access rules are ready for a full monthly rollout.
€431 one-time. A bounded audit for up to 25 representative files (maximum 200 pages total) and up to 15 fixed questions, with citation checks and a go/no-go report before recurring Team RAG.
Upgrade path: when the recurring Team RAG scope is approved within 30 days, the €431 audit fee is credited against its setup fee and confirmed before provisioning. Extra files, pages, questions, or rework are quoted separately.
We start with a limited set of real documents, screenshots, policies, tickets, or manuals and turn them into a fixed retrieval benchmark instead of ingesting everything blindly.
Qwen3-Embedding 0.6B/4B and Qwen3-Reranker 0.6B/4B are the default audit candidates. The 8B path stays optional when corpus size, latency, and the 20 GB VRAM budget justify it.
The result is a short report with top-k misses, citation quality, reranker gains, storage notes, privacy boundaries, and a clear decision: fix sources, run a pilot, or move to Team RAG.
Current local evidence (2026-08-02): qwen3-embedding:0.6b + gpt-oss:20b passed a fixed public EZOS Team RAG/agent benchmark: 5 retrieval cases, 0.8 top-1 retrieval accuracy, and 3/3 grounded answer plus agent-boundary cases. This proves the benchmark path on public sample data, not customer-corpus production performance.
Our promise: this audit gives you real evidence of search quality and a clear next step. It makes no live-performance claim until the GPU, Ollama, and your chosen models pass tests on your own server.
Use this page to pick the right starting point, then confirm the exact setup, renewal terms, and any test notes at checkout.
from $295.10/mo
Already have your own GPU server or rented one? We manage Ollama, Open WebUI, updates, and monitoring on it. (BYO means bring your own server.)
from $699.42/mo
For private local model hosting with a managed operating layer and no third-party AI API required by default.
from $999.60/mo
For teams that need document-assisted local AI, curated embeddings, user roles, and clearer knowledge workflows.
from $1,499.90/mo
For production business use where security, audit preparation, update windows, and support scope matter.
Before we make any live-AI claim, we check the GPU drivers and confirm Ollama is healthy. A 20 GB RTX 4000 Ada card is great for small-to-medium models, not the very largest frontier models.
Public starting prices are shown in USD and were checked against the secure cart on 2026-08-03. Checkout defaults to EUR, can switch currency, and remains authoritative for setup fees, term, renewal, taxes, and final total.