XARIV
← Products

Self-hosted open-weight serving

XARIV Relay

Download open-weight models, serve them on your own Mac with llama.cpp (or vLLM on Linux NVIDIA), and compare two local endpoints side by side — TTFT, stream speed, and output.

Capabilities

  • —Hugging Face GGUF download or local import
  • —Deploy to localhost (llama.cpp Metal on Mac)
  • —OpenAI-compatible /v1 endpoints
  • —Split-pane live compare of two candidates

Status: preview