Documentation
Run it, route it, read the numbers.
Everything you need to put Marbor in front of your cluster -- connect your tools, ship to production, and understand exactly what it's saving you.
Guides · 5 pages
Integrations
marbor exposes an Ollama-compatible API on port 11434 and passes through Ollama's OpenAI-compatible /v1 endpoints unchan...
Read →Known limitations
This page documents what marbor does not do, what has been tested, and what to plan around in production. Infrastructure...
Read →Savings math
This document is the financial model behind marbor's savings tracking. It is designed for infrastructure directors and f...
Read →Use cases
The infrastructure control plane Ollama doesn't ship: secure multi-tenant access, hardware-aware load balancing, cost-aw...
Read →Backup & Restore
marbor is DB-first: marbor.db SQLite holds nodes, API keys, routing rules, warm-state...
Read →Deployment · 4 pages
Production
This guide covers running marbor in a real production environment. There is no config file - marbor is DB-first marbor.d...
Read →AWS EC2
Run one marbor endpoint in front of one or more Ollama GPU boxes on EC2....
Read →GPU node registration
Register many GPU nodes' runtime endpoints with marbor at once, without clicking...
Read →marbor agent enrollment
Enroll and install the marbor agent on many already-registered GPU nodes at once,...
Read →Integrations · 4 pages
Continue
Continue is an open-source AI coding assistant. Point it at marbor and every completion request routes through warm-firs...
Read →LibreChat
LibreChat supports Ollama as a custom endpoint. Point it at marbor to get warm-first routing across multiple GPU nodes, ...
Read →LiteLLM
LiteLLM is a popular and powerful API gateway for provider abstraction, enterprise authentication, and user-level rate l...
Read →Open WebUI
Point Open WebUI at marbor instead of a single Ollama box and you get warm-first routing across all your GPU nodes, cost...
Read →