Your AI, your data, sovereign without the cloud.
KNUT unites existing hardware, from NVIDIA to Apple Silicon, across your local network into a single interface. OpenAI and Anthropic compatible, hosted on-premise, with zero variable token pricing.
Why companies rely on local & private AI infrastructure.
The productive adoption of artificial intelligence unfolds its full potential when trade secrets remain protected and ongoing operating costs stay transparent and predictable.
Confidentiality & Professional Secrecy
Sensitive documents, client contracts, patient records, or proprietary source code demand utmost discretion. With local execution, all inputs stay strictly within your company network—with zero third-party transmission.
Predictable Operating Costs as Usage Grows
Under intensive workloads, automated n8n pipelines, or continuous RAG queries, fixed operating costs based on actual electricity usage provide far greater predictability than volatile volume-based pricing.
Combine Existing Machines Seamlessly
Most businesses already possess capable desktop PCs, workstations, or Macs. KNUT connects these machines across your local network into a shared pool, allowing demanding frontier models to run directly on-premise.
Pool your hardware into a private AI compute cluster.
Acting as a resilient network knot, KNUT connects commercial GPUs and computers (such as NVIDIA RTX models or Apple Silicon Macs clustered with PCs) over your local network. Businesses can run modern language models for internal workflows, confidential documents, and automations entirely on-site—pragmatically, with zero data leaks, leveraging the hardware already on your desks.
Exemplary data based on pure GPU power consumption under inference load (excluding host idle power): the lowest-cost measured configuration in the KNUT cluster is Google Gemma 4 Suite (MoE-26B-A4B-Q4) at 115.9 t/s and 180 W GPU load, consuming 0.43 kWh per 1M tokens. The actual operating cost depends on your electricity contract: €0.04 with self-generated solar (€0.10/kWh), €0.10 under commercial SME contracts (€0.24/kWh), €0.15 under standard grid power (€0.36/kWh). Larger models and higher precision quantizations scale accordingly. Full benchmark measurements are detailed below. Hardware acquisition and depreciation are not included.
Three strands for the perfect knot
KNUT stands for the Nordic knot: just like in seafaring, it binds individual strands into a resilient unit. In your LAN, KNUT bundles your machines—whether NVIDIA or Apple Silicon—into a unified pool, powered by llama.cpp with custom distributed orchestration and telemetry.
Sovereignty
Data stays inside your network
Documents, legal files, and patient records are processed locally inside your LAN without third-party processors. The foundation for full compliance and GDPR peace of mind.
Pooling
Shared video memory across the network
KNUT pools NVIDIA cards and Macs into a shared VRAM pool: two RTX 5070 Tis and two Mac Mini M4s combine for ~80 GB VRAM. What cannot fit on a single device runs smoothly across the cluster (from ~€0.11 / 1M tokens*).
Integration
Single connection via Port 9600
All endpoints converge in one place: a drop-in API compatible with OpenAI and Anthropic standards. n8n pipelines, Open WebUI, Cursor, and custom Python agents connect instantly.
Exemplary data based on pure GPU power consumption under inference load (excluding host idle power): the lowest-cost measured configuration in the KNUT cluster is Google Gemma 4 Suite (MoE-26B-A4B-Q4) at 115.9 t/s and 180 W GPU load, consuming 0.43 kWh per 1M tokens. The actual operating cost depends on your electricity contract: €0.04 with self-generated solar (€0.10/kWh), €0.10 under commercial SME contracts (€0.24/kWh), €0.15 under standard grid power (€0.36/kWh). Larger models and higher precision quantizations scale accordingly. Full benchmark measurements are detailed below. Hardware acquisition and depreciation are not included.
Distributed inference,
measured in practice.
KNUT pools GPUs and Apple Silicon* across your LAN. Benchmarked on a reference cluster spanning three CUDA generations, showing exact real-world throughput and power metrics.
Qwen3.8-27B
Cutting-edge dense architecture with State-Space Modeling (SSM) and Multi-Token Prediction (MTP). Delivers exceptional reasoning and coding capabilities across a 256k context window.
Measured Quantization Levels (Qwen3.8-27B)
Reference cluster: up to 3 CUDA generations (RTX 5070 Ti + 4060 Ti + 3060) with up to 44 GB VRAM| Quantization | File Size | Throughput | GPU Power (W) | GPU Cost / 1M* |
|---|---|---|---|---|
| UD-IQ2_XXS | 9.0 GB | 65.7 t/s | 348 W | 0.53 € |
| UD-IQ2_M | 10.3 GB | 58.4 t/s | 330 W | 0.56 € |
| UD-Q3_K_XL | 13.4 GB | 43.3 t/s | 276 W | 0.63 € |
| UD-Q4_K_XL | 17.9 GB | 38.2 t/s | 245 W | 0.64 € |
| UD-Q6_K_XL | 25.9 GB | 27.5 t/s | 232 W | 0.84 € |
| UD-Q8_K_XL | 31.5 GB | 19.2 t/s | 227 W | 0.87 € |
Measured on a reference cluster across three CUDA generations: 1× RTX 5070 Ti (node nexus), 1× RTX 4060 Ti (node vector), and 1× RTX 3060 12GB (node wall-e), as well as Apple Silicon. While this setup does not match high-end datacenter cards in sheer memory bandwidth, the three cards combine 44 GB VRAM from affordable mainstream hardware. Models up to 35B and beyond run smoothly across the LAN without hitting individual VRAM limits. Measurements reflect pure GPU power draw under inference load at an exemplary €0.36/kWh.
Cloud vs. Self-Hosted.
A model calculation based on real measurements: choose one of the regional reference tariffs (based on Northern Germany) or customize the electricity price to match your exact contract—from solar PV to commercial tariffs.
Standard retail electricity price for offices and practices including all grid fees and taxes.
Model calculation, no guarantee. Assumes a cluster throughput of 100 t/s, 700 W combined GPU power draw under load, and 140 kWh monthly idle baseline (24/7 node operation). 1 million tokens correspond roughly to 1,600 standard DIN A4 text pages (~600 tokens per page including input and output; a standard 8 cm lever-arch folder holds ~500 sheets). Actual consumption depends on model architecture, quantization, and workload. Cloud comparison uses midpoint estimates across three model classes, not list prices of specific vendors. Hardware purchase and depreciation costs are not included.
KNUT in Numbers
- External API Costs
- 0 €for on-premise operation in your LAN
- Electricity from
- ~0,11 €per 1M tokens*
- Data Sovereignty
- 100 %of requests processed on your network
- Context Window
- 256k+model dependent, scales with the pool
Exemplary data based on pure GPU power consumption under inference load (excluding host idle power): the lowest-cost measured configuration in the KNUT cluster is Google Gemma 4 Suite (MoE-26B-A4B-Q4) at 115.9 t/s and 180 W GPU load, consuming 0.43 kWh per 1M tokens. The actual operating cost depends on your electricity contract: €0.04 with self-generated solar (€0.10/kWh), €0.10 under commercial SME contracts (€0.24/kWh), €0.15 under standard grid power (€0.36/kWh). Larger models and higher precision quantizations scale accordingly. Full benchmark measurements are detailed below. Hardware acquisition and depreciation are not included.
NVIDIA and Apple Silicon Clustered Together.
From a Mac Mini or older RTX 3060 to the latest RTX 5070 Ti: KNUT bundles generations and architectures into a shared VRAM pool. Any office or lab computer can flexibly act as master or worker in the cluster.
Drop-in API Interfaces
Port 9600 (FastAPI Cluster Proxy)
SSE streaming inference, function calling / tool use, and JSON schema validation.
Translation of Claude tools and tool results into llama-server calls.
Guaranteed output formatting in predefined JSON schemas for workflows.
Live VRAM, wattage, GPU temperature, and token throughput via SSE & REST.
Seamless Tool Integration
Standard Interfaces: Simply swap the endpoint URL
Connection via default OpenAI node, usually zero code changes required.
Full-featured ChatGPT interface running in your LAN with user management.
Local coding assistant with large context window & MTP speculative speedup.
RAG pipelines for confidential documents strictly within your network.
For organizations where confidentiality and cost efficiency come first.
KNUT was built for businesses seeking to deploy modern AI models productively without relinquishing control of confidential data to external platforms.
Law Firms & Notaries
Processing Confidential Briefs with Zero Third-Party Risk
Contract reviews, case summaries, and client correspondence are processed inside your office network. No third-party cloud transfers, eliminating data processing compliance burdens.
Medical Practices, Labs & Clinics
On-Premise Medical Reports & Documentation
Analyze medical letters, lab diagnostics, and literature on private hardware. Provides a safe foundation for handling special categories of personal data under GDPR Art. 9.
Digital & Marketing Agencies
Unlimited Compute for Content, Research & Automation
Keep workflows running smoothly across content generation, automated research, or agentic experiments. Predictable operating costs by pooling existing agency hardware.
Software & Tech Teams
Local Coding Models with Large Context for Cursor & Cline
Intellectual property and proprietary codebases remain strictly within your development environment. Fast completions and large refactors with zero data leaks.
Ready for your own,
private AI infrastructure?
Two minutes, three fields. Be the first to know when KNUT is ready for your business and what your existing hardware can deliver.