Step 2 · Benchmark
Inference Benchmarking & Telemetry
See the latency profile before you serve a single request.
Point Pulse at a public dataset like ShareGPT or paste your own prompts. It replays the request mix on your target GPU and summarizes TTFT, ITL, TPOT, end-to-end latency, throughput, and live-style GPU telemetry.
Run real benchmarks on your machine
Queue jobs in Workspace → connect xariv-pulse CLI → track results in your profile table.