XARIV
← Products

Self-hosted open-weight serving

XARIV Relay

Download open-weight models, serve them on your own Mac with llama.cpp (or vLLM on Linux NVIDIA), and compare two local endpoints side by side — TTFT, stream speed, and output.

Capabilities

  • Hugging Face GGUF download or local import
  • Deploy to localhost (llama.cpp Metal on Mac)
  • OpenAI-compatible /v1 endpoints
  • Split-pane live compare of two candidates

Status: preview