← Products
Self-hosted open-weight serving
XARIV Relay
Download open-weight models, serve them on your own Mac with llama.cpp (or vLLM on Linux NVIDIA), and compare two local endpoints side by side — TTFT, stream speed, and output.
Capabilities
- —Hugging Face GGUF download or local import
- —Deploy to localhost (llama.cpp Metal on Mac)
- —OpenAI-compatible /v1 endpoints
- —Split-pane live compare of two candidates
Status: preview