Need real llama.cpp endpoints? Deploy locally after models are cached.

XARIV Relay · Compare

Same prompt. Two local candidates.

Llama 3.2 1B — F16 vs Q4. This page replays a typical Mac Metal timing profile so you can show the speed gap without downloading weights on Vercel.

Illustrative replay (not live GPU inference in this browser). Edit the prompt below — both panes still stream the same canned answer so the speed gap stays clear. Real models cache in ~/.xariv/relay/models/.

Left · baseline

Llama 3.2 1B Instruct · F16 (non-quantized)

localhost:8090

Right · quantized

Llama 3.2 1B Instruct · Q4_K_M (quantized)

localhost:8091

Llama 3.2 1B Instruct · F16 (non-quantized)

localhost:8090

TTFT

tok/s

est. tokens

Output streams here…

Llama 3.2 1B Instruct · Q4_K_M (quantized)

localhost:8091

TTFT

tok/s

est. tokens

Output streams here…